Long text processing method and device, equipment and medium
By obtaining the keyword sequence and word weight information of long texts for feature mapping, the problems of efficiency and information integrity in existing long text representation methods are solved, efficient and accurate long text representation is achieved, and the effects of text retrieval and matching are improved.
Patent Information
- Application Number
- CN202410392175.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-02
- Publication Date
- 2025-10-14
AI Technical Summary
Existing long text representation methods have shortcomings in efficiency and information completeness. The truncation method leads to information loss, and the sharding and pooling method ignores text connections and is inefficient.
By obtaining the keyword sequence and word weight information of the target long text, feature mapping is performed to obtain text mapping features. Feature extraction is performed based on this feature to improve the integrity and continuity of text representation, and word weights are introduced to express the importance of keywords.
It achieves efficient modeling for long text processing, improves the expression accuracy and specificity of text content, and significantly improves the recall rate and accuracy of text retrieval and matching.
Smart Images

Figure CN120781801A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a long text processing method, device, equipment and medium. BACKGROUND
[0002] Feature representation of long text is the data basis of various text data application scenarios, such as similarity comparison element for long text retrieval or matching tasks, and implementation of text recall and precision ranking. Existing long text representation extraction methods mainly include truncation method and slicing pooling method. The truncation method usually concatenates the text of the beginning and end of the article for tokenization and input into the text model, or first summarizes the entire article by means of an abstract generation model, and then extracts the representation, so as to ensure that the length of the text input into the text model does not exceed the maximum limit. The slicing pooling method retains the entire article content, and divides the article into text slices (batch) by sliding window, and then inputs the text model (such as BERT) in sequence, and then pools the features of each slice to obtain the text representation. However, the truncation method is fast, but it causes information loss and uneven input content, and the representation effect depends on the pre-posed abstract generation model. The slicing pooling method does not have direct information loss, but it ignores the connection between different segments, essentially does not model the entire article, and at the same time, the reasoning cost increases linearly with the length of the article. Figure 3 SUMMARY
[0003] The present application provides a long text processing method, device, equipment and medium, which can significantly improve the modeling efficiency and text representation effect of long text processing.
[0004] In one aspect, the present application provides a long text processing method, which comprises:
[0005] obtaining a keyword sequence and word weight information of a target long text, the keyword sequence comprising a plurality of keywords corresponding to the target long text and being ordered based on the text order of the target long text, and the word weight information comprising a word weight of each keyword in the plurality of keywords, the word weight being used to represent the importance of the keyword to the target long text;
[0006] performing feature mapping on the keyword sequence and the word weight information to obtain a text mapping feature, the text mapping feature being used to represent the word feature, weight feature and position feature of the keyword in the keyword sequence;
[0007] performing feature extraction on the target long text based on the text mapping feature to obtain a long text representation result.
[0008] In another aspect, a long text processing device is provided, which comprises:
[0009] An acquisition module is configured to acquire a keyword sequence and word weight information of a target long text, wherein the keyword sequence includes a plurality of keywords corresponding to the target long text and is sorted based on the text order of the target long text; the word weight information includes a word weight of each keyword in the plurality of keywords, and the word weight is used to represent the importance of the keyword to the target long text;
[0010] Feature mapping module: used to perform feature mapping on the keyword sequence and the word weight information to obtain text mapping features, wherein the text mapping features are used to characterize the word features, weight features and position features of the keywords in the keyword sequence;
[0011] Feature extraction module: used to extract features of the target long text based on the text mapping features to obtain a long text representation result.
[0012] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the long text processing method as described above.
[0013] On the other hand, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction or at least one program segment, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the long text processing method as described above.
[0014] On the other hand, a server is provided, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the long text processing method as described above.
[0015] On the other hand, a terminal is provided, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the long text processing method as described above.
[0016] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program comprises computer instructions. When the computer instructions are executed by a processor, the long text processing method as described above is implemented.
[0017] The long text processing method, apparatus, device, storage medium, server, terminal, computer program, and computer program product provided in this application have the following technical effects:
[0018] The application first acquires a keyword sequence and word weight information of a target long text, performs feature mapping on the keyword sequence and the word weight information to obtain text mapping features, and then performs feature extraction on the target long text based on the text mapping features to obtain a long text representation result. The keyword sequence includes a plurality of keywords corresponding to the target long text and is sorted based on the text order of the target long text. The word weight information includes a word weight of each keyword in the plurality of keywords, which can represent the importance of the keyword to the target long text. The corresponding text mapping features are used to represent the word features, weight features and position features of the keywords in the keyword sequence. The global core content is grasped through the extraction of the keywords and the sorting of the text sequence, ensuring the integrity, balance and continuity of the expression of the key content of the long text, breaking away from the length limitation of the text, and improving the text content modeling effect and efficiency. Meanwhile, the word weight information is introduced to express the importance of the keywords, effectively improving the expression accuracy and specificity of the text content on the premise of efficient representation extraction of the long text, and significantly improving the effective content recall rate and accuracy in the text retrieval and matching scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0020] Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application;
[0021] Figure 2 is a flowchart of a long text processing method provided by an embodiment of the present application;
[0022] Figure 3 is a principle diagram of a text processing method provided by the prior art;
[0023] Figure 4 is a flowchart of another long text processing method provided by an embodiment of the present application;
[0024] Figure 5 is a flowchart of another long text processing method provided by an embodiment of the present application;
[0025] Figure 6 is a principle diagram of a long text processing method provided by an embodiment of the present application;
[0026] Figure 7 is a principle framework diagram of a long text processing method provided by an embodiment of the present application;
[0027] Figure 8 Another long text processing method provided by the embodiment of the present application has a principle framework diagram as shown in FIG. 6.
[0028] Figure 9 A structural framework diagram of a feature extraction module provided by the embodiment of the present application has a structural framework diagram as shown in FIG. 7.
[0029] Figure 10 Another structural framework diagram of a feature extraction module provided by the embodiment of the present application has a structural framework diagram as shown in FIG. 8.
[0030] Figure 11 A flow framework diagram of a model training method provided by the embodiment of the present application has a flow framework diagram as shown in FIG. 9.
[0031] Figure 12 A principle framework diagram of a model training method provided by the embodiment of the present application has a principle framework diagram as shown in FIG. 10.
[0032] Figure 13 Another principle framework diagram of a model training method provided by the embodiment of the present application has a principle framework diagram as shown in FIG. 11.
[0033] Figure 14 A framework schematic diagram of a long text processing device provided by the embodiment of the present application has a framework schematic diagram as shown in FIG. 12.
[0034] Figure 15 A hardware structure block diagram of an electronic device for executing a long text processing method provided by the embodiment of the present application has a hardware structure block diagram as shown in FIG. 13. DETAILED DESCRIPTION
[0035] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0036] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or sub-modules does not necessarily have to be limited to those steps or sub-modules clearly listed, but can include other steps or sub-modules not clearly listed or inherent to these processes, methods, products or devices.
[0037] Before the embodiments of the present application are further described in detail, the terms and names involved in the embodiments of the present application are explained, and the terms and names involved in the embodiments of the present application are applicable to the following explanations.
[0038] Artificial Intelligence (AI) is the theory, method, technology and application system of using digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain optimal results. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0039] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-training model technology, operation / interaction system, mechatronics, etc. Among them, the pre-training model is also called large model, basic model, which can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning, etc.
[0040] Nature Language processing (NLP) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between people and computers using natural language. Natural language processing involves natural language, i.e. the language used in daily life, and is closely related to linguistics, as well as computer science and mathematics. The pre-training model, an important technology for model training in the field of artificial intelligence, is developed from the Large Language Model (LLM) in the field of NLP. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question and answer, knowledge graph, etc.
[0041] A pre-training model, also known as a cornerstone model or a large model, refers to a deep neural network (DNN) with large parameters. It is trained on massive amounts of unlabeled data. Leveraging the function approximation capabilities of large-parameter DNNs, the pretrained machine learning (PTM) extracts common features from the data. Through techniques such as fine tuning, efficient parameter fine tuning (PEFT), and prompt-tuning, it is then adapted for downstream tasks. Therefore, pre-trained models can achieve ideal results in few-shot or zero-shot scenarios. Based on the data modality processed, PTMs can be categorized into language models (ELMO, BERT, GPT), vision models (swin-transformer, ViT, V-MOE), speech models (VALL-E), and multimodal models (ViBERT, CLIP, Flamingo, Gato). Multimodal models are those that represent features from two or more data modalities. Pre-trained models are important tools for outputting artificial intelligence generated content (AIGC) and can also serve as a universal interface for connecting multiple task-specific models.
[0042] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.
[0043] BERT: Bidirectional Encoder Representations from Transformers, a bidirectional encoder based on the Transformer structure, is a deep learning network structure commonly used in the field of natural language processing.
[0044] TF-IDF: Term Frequency-Inverse Document Frequency, is a text weighting technology based on term frequency and inverse document frequency, commonly used in information retrieval and data mining.
[0045] RNN: Recurrent Neural Network, a recurrent neural network, is a common neural network.
[0046] In recent years, with the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0047] See also Figure 1 , Figure 1 02 is a schematic diagram of an application environment provided in an embodiment of the present application, which may include at least a server 02. The server 02 in the embodiment of the present application may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0048] Specifically, cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to enable data computing, storage, processing, and sharing. Cloud technology can be applied in a variety of fields, such as healthcare cloud, cloud IoT, cloud security, cloud education, cloud conferencing, artificial intelligence cloud services, cloud applications, cloud calling, and cloud social networking. Based on the cloud computing business model, cloud technology distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network that provides resources is called the "cloud." To users, the resources in the cloud appear infinitely scalable and can be accessed at any time, used on demand, and expanded at any time, with a pay-per-use policy. Providers of cloud computing infrastructure establish a cloud computing resource pool (referred to as a cloud platform, commonly referred to as IaaS (Infrastructure as a Service)) and deploy various types of virtual resources within the resource pool for external clients to choose from. The cloud computing resource pool primarily includes computing devices (virtualized machines, including operating systems), storage devices, and network devices.
[0049] Specifically, the server 02 mentioned above may include a physical device, which may specifically include a network communication submodule, a processor, a memory, etc., and may also include software running in the physical device, which may specifically include an application program, etc.
[0050] Specifically, server 02 is used to obtain the keyword sequence and word weight information of the target long text, and perform feature mapping on the keyword sequence and word weight information to obtain text mapping features. The text mapping features are used to characterize the word features, weight features and position features of the keywords in the keyword sequence; then, feature extraction of the target long text is performed based on the text mapping features to obtain a long text characterization result.
[0051] like Figure 1 As shown, the application environment may also include a terminal 01, and the terminal 01 and the server 02 may be directly or indirectly connected via wired or wireless communication, which is not limited in this application. Specifically, the terminal 01 may include physical devices such as smart phones, desktop computers, tablet computers, laptops, digital assistants, augmented reality (AR) / virtual reality (VR) devices, intelligent voice interaction devices, smart home appliances, smart wearable devices, and vehicle-mounted terminal devices, and may also include software running in physical devices, such as applications. In an embodiment of the present application, the terminal 01 may be used to send a long text to be processed to the server 02, so that the server 02 performs the above-mentioned text processing, and then receives recall text information based on the long text characterization results.
[0052] Furthermore, it is understandable that Figure 1 What is shown is merely an application environment of a long text processing method. The application environment may include more or fewer nodes, and this application does not impose any limitation thereto.
[0053] The application environment involved in the embodiments of this application, or the terminal 01 and server 02 in the application environment, can be a distributed system formed by connecting a client and multiple nodes (any form of computing device connected to the network, such as a server or user terminal) through network communication. The distributed system can be a blockchain system that can provide the above-mentioned long text processing services, model training services, and related data storage services.
[0054] The following is an introduction to the technical solution of this application based on the above application environment. The embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc. Please refer to Figure 2 , Figure 2It is a flowchart of a long text processing method provided by an embodiment of the present application. This specification provides method operation steps such as the embodiment or flowchart, but may include more or fewer operation steps based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the order of many steps and does not represent the only execution order. When the actual system or server product is executed, it can be executed in sequence or in parallel according to the method shown in the embodiment or the accompanying drawings (for example, in a parallel processor or multi-threaded processing environment). Specifically, as Figure 2 As shown, the method may include the following steps S201-S205:
[0055] S201: Obtain keyword sequence and word weight information of the target long text.
[0056] Specifically, the target long text can be from any content category, such as fiction, news, or scientific text, and its length type is long text. Optionally, the long text type can be text with a length greater than or equal to a text length threshold. This text length threshold can represent the upper limit of the number of word segments or the upper limit of the number of words in the text. The text length threshold can be set based on actual business needs and scenarios.
[0057] In some embodiments, before S201 , the method further includes: obtaining an original long text; and performing text cleaning and word segmentation processing on the original long text to obtain a target long text.
[0058] Specifically, original long text refers to unprocessed document text, such as original novel paragraph documents or original news text, and text cleaning refers to the removal of text cleaning. Text cleaning refers to modifying the text by deleting useless information, fixing errors and noise, etc. to obtain clean text. Text cleaning preprocessing can filter out meaningless characters in the original text or content that interferes with the modeling of text content, including but not limited to URLs, punctuation marks, emoticons, very long numbers, etc. The text obtained after text cleaning is a text that can be segmented.
[0059] Specifically, word segmentation involves segmenting the cleaned text based on a preset dictionary, splitting the continuous text into multiple text segments. The resulting target long text is equivalent to the collection of the previously obtained text segments, sorted in text order. Through text cleaning and word segmentation, we obtain a target long text suitable for keyword extraction, remove text noise interference, and improve the accuracy of keyword extraction and word weight information acquisition.
[0060] There are often certain differences in the distribution of vocabulary in different application scenarios. The preset dictionary can be set based on the application scenario, or the corresponding category dictionary can be selected according to the category of the original long text. For example, scientific articles can use a custom dictionary with professional vocabulary, or ordinary novels or news documents can use a regular dictionary. In this way, the use of an adapted dictionary can ensure the reliability and integrity of subsequent text modeling. For example, through the use of a custom dictionary, it can ensure that common vocabulary or professional vocabulary in subsequent application scenarios are not segmented and can enter the subsequent processing stage as a complete unit.
[0061] Specifically, the keyword sequence includes multiple keywords corresponding to the target long text and is sorted based on the text order of the target long text. Keywords refer to text segmentations that can serve as indexes for the target long text and are extracted from the text segmentation set of the target long text. The keyword sequence is a sequence formed by splicing multiple extracted keywords. Text order sorting based on the target long text means sorting multiple keywords in the order of text reading to preserve the semantic and textual connections between the content represented by the keywords as much as possible, which is conducive to the textual expression accuracy of the representation results.
[0062] Specifically, the word weight information includes the word weight of each keyword in the multiple keywords. The word weight is used to represent the importance of the keyword to the target long text. A higher word weight indicates a higher importance, and vice versa. The data format of the word weight information can be a concatenation of the word weights of each keyword in the keyword sequence based on the order of the keyword sequence, which is equivalent to a word weight vector.
[0063] In some embodiments, reference Figure 4 , S201 may include:
[0064] S2011: extract keywords from the target long text to obtain multiple keywords;
[0065] S2012: Sort and concatenate multiple keywords based on the text order of the target long text to obtain a keyword sequence;
[0066] S2013: Determine the word weight of each keyword based on the category distinction information of the keyword and the frequency of occurrence of the keyword in the target long text to obtain word weight information.
[0067] Specifically, the keyword extraction can be performed using an existing method capable of extracting text keywords. The extracted keywords carry the key content information of the original long text and are valid text segmentation words with practical meaning. Multiple keywords can cover the subject content of each paragraph in the full text, achieving global coverage of the text information. Valid text segmentation words refer to words with text meanings. For example, auxiliary words, conjunctions, and personal pronouns are all invalid text segmentation words. Optionally, sorting and splicing multiple keywords based on text order means using the first occurrence position of the keyword in the target long text as the sorting sequence number, or using the average position of the keyword in each occurrence position in the target long text as the sorting sequence number, or the closer the keyword position is, the closer it is to the front of the keyword sequence, and vice versa. In some cases, the keywords with text connections in the target long text can be connected to obtain keyword fragments, and then the obtained keyword fragments and other keywords are sorted based on text order. The sequence number of the keyword fragment can be the text position of the first occurrence of the keyword fragment, or the average position of the keyword fragment's occurrence positions.
[0068] Specifically, the frequency of occurrence refers to the ratio of the number of times a keyword appears in the target long text to the total number of text segmentations of the target long text. The higher the frequency of occurrence of the keyword, the higher the word weight, and vice versa. Category differentiation information is used to indicate the text category differentiation ability of the keyword. The higher the value corresponding to the category differentiation information, the greater the differentiation ability and the higher the word weight. Conversely, the smaller the differentiation ability and the lower the word weight. The higher the proportion of documents containing a keyword in the corpus to the total number of documents, the lower the value corresponding to the category differentiation information. Conversely, the higher the value, that is, the smaller the value, the better the document specificity and discrimination of the keyword. The corpus can be obtained based on actual task requirements, or it can be a basic corpus. In some embodiments, the word weight can be the product of the value of the category differentiation information and the frequency of occurrence. For example, the frequency of occurrence can be the word frequency (TF) of the keyword, and the category differentiation information can be the inverse document frequency (IDF) of the keyword. The word weight is the product of TF and IDF.
[0069] In some embodiments, the keyword sequence includes all keywords extracted in S2011 to ensure the integrity of the text expression. In other embodiments, due to the input limitations of the model, a preset number of words (N) needs to be set. When the number of extracted keywords exceeds the preset number of words, the keywords can be sorted from large to small based on their word weights. The top N keywords in the sorting can be used as target keywords to perform keyword concatenation in S2012 to obtain a keyword sequence. N can be, for example, 100.
[0070] By extracting keywords and sorting text in sequence, we generate keyword sequences that can express the key content information of long texts and semantically coherent keywords, thereby improving the completeness and accuracy of the semantic information of subsequent text representations. In addition, we determine word weights from two dimensions through frequency of occurrence and category differentiation information, so that the word weight information can reflect the importance of the keyword in the text while expressing the text differentiation of the keyword, thus bringing the representation results of similar texts closer and the representation of different types of texts further apart.
[0071] S203: Perform feature mapping on the keyword sequence and word weight information to obtain text mapping features.
[0072] Specifically, feature mapping is used to vectorize the keyword sequence and word weight information, resulting in text mapping features in the form of feature vectors. Text mapping features are used to characterize the word features, weight features, and position features of the keywords in the keyword sequence, thereby expressing word meaning, word importance, and position in the text.
[0073] In practical applications, word features are word vectors obtained by embedding keywords, weight features are weight vectors obtained by embedding word weights, and position features are position vectors obtained by embedding the position information of keywords in the keyword sequence. The above feature embedding is implemented using a target text model, which includes an embedding module and a feature encoding module. The embedding module implements the mapping of the input keyword sequence and word weight information to the feature space. Accordingly, in some embodiments, reference Figure 5 , S203 includes S2031-S2034:
[0074] S2031: performing word feature embedding on each keyword in the keyword sequence based on the embedding module of the target text model to obtain word sequence features;
[0075] S2032: Performing position feature embedding on the sequence position of each keyword in the keyword sequence based on the embedding module to obtain a position sequence feature;
[0076] S2033: Perform weight feature embedding on each word weight in the word weight information based on the embedding module to obtain a weight sequence feature;
[0077] S2034: Fusing word sequence features, position sequence features, and weight sequence features to obtain text mapping features.
[0078] Specifically, the embedding module sets a trainable first parameter matrix, a second parameter matrix and a third parameter matrix. First, the keyword sequence is input into the embedding module, each keyword is mapped to a word segmentation coding ID (token ID), and feature embedding is performed on each word segmentation coding ID based on the first parameter matrix to obtain a word segmentation feature (word embedding). Then, word sequence features are obtained after splicing the keyword sequence according to the order, and the word sequence features are vectorized representations of the keyword sequence. Exemplarily, keywords can be mapped to corresponding word segmentation coding IDs by, for example, BERT tokenizer. At the same time, a weight feature (initial weight embedding) is assigned to the word weight of each keyword. The weight feature is obtained by weighted embedding of the word weight by the second parameter matrix. After splicing the weight features according to the order of the keyword sequence, a weight sequence feature is obtained. The weight sequence feature is a vectorized representation of the word weight information. It can be understood that the word weight information is obtained after weight normalization processing of each keyword in the keyword sequence. In addition, each keyword has a corresponding position feature (position embedding) to retain the relative position relationship of the extracted keyword group. After splicing the position features according to the order of the keyword sequence, a position sequence feature is obtained.
[0079] Specifically, fusion here refers to the summation of word sequence features, position sequence features, and weight sequence features to obtain text mapping features. Through word embedding, position embedding, and weight embedding, we obtain text embedding features that can comprehensively express word information, word text order information, and importance information, which is beneficial to the accuracy of text representation.
[0080] S205: Extract features of the target long text based on the text mapping features to obtain a long text representation result.
[0081] In summary, the above technical solution grasps the global core content through keyword extraction and text sequence sorting, ensures the integrity, balance and continuity of the expression of key content in long texts, gets rid of the text length limit, and improves the text content modeling effect and efficiency; at the same time, it introduces word weight information to express the importance of keywords, and effectively improves the expression accuracy and specific discrimination of text content under the premise of realizing efficient representation extraction of long texts, and significantly improves the effective content recall rate and accuracy in scenarios such as long text retrieval and long text matching.
[0082] In an embodiment of the present application, feature extraction is implemented through a feature encoding module of the target text model. Accordingly, in some embodiments, S205 can be specifically as follows: inputting text mapping features into the feature encoding module of the target text model for feature encoding based on the attention mechanism to obtain a long text representation result.
[0083] Specifically, the feature encoding module is built on a transformer and employs an attention mechanism for feature encoding. During the feature encoding process, text embedding features are cross-referenced to fully leverage contextual information, efficiently obtaining a comprehensive and accurate long text representation. This long text representation is a vector representation of the target long text. Specifically, the long text representation can be obtained by pooling the feature vectors output by the hidden layer of the feature encoding module.
[0084] Accordingly, refer to Figure 6 In one embodiment, the long text processing method may specifically include: S11, obtaining the original long text; S12, performing text cleaning and word segmentation processing on the original long text to obtain the target long text; S13, inputting the target long text into the keyword extraction module to extract keywords, obtaining multiple keywords and the word weight of each keyword, so as to generate word weight information; S14, sorting and splicing the multiple keywords based on the text order of the target long text to obtain a keyword sequence; S15, performing word feature embedding on each keyword in the keyword sequence based on the embedding module of the target text model to obtain word sequence features; S16, embedding the keyword sequence based on the embedding module The position feature of the sequence position of each keyword in the word weight information is embedded to obtain the position sequence feature; S17, based on the embedding module, the weight feature of each word in the word weight information is embedded to obtain the weight sequence feature; S18, the word sequence feature, the position sequence feature and the weight sequence feature are added to obtain the text mapping feature; S19, the text mapping feature is input into the feature encoding module for feature extraction, and the feature vector of the hidden state output by the hidden layer of the feature encoding module is pooled to obtain the long text representation result. Compared with the text feature modeling method based on RNN network, the long text processing method of the present application has a more accurate and comprehensive representation effect. In some embodiments, the target text model can be constructed based on a model similar to BERT.
[0085] In other embodiments, reference Figure 7 , S205 may include S301-S305:
[0086] S301: sorting multiple keywords of the target long text by weight based on word weight information to obtain a reference sequence.
[0087] Specifically, weight sorting refers to sorting keywords according to the value of word weight from large to small or from small to large, preferably from large to small, thereby reflecting the importance sorting information of keywords in the reference sequence.
[0088] S303: Perform feature mapping on the reference sequence based on the embedding module of the target text model to obtain reference mapping features.
[0089] Specifically, the reference mapping feature is used to characterize the word feature and importance position feature of the keyword in the reference sequence, so as to carry the importance of the keyword to the long text while expressing the semantic information. In a specific embodiment, S303 may include:
[0090] S3031: Perform word feature embedding on each keyword in the reference sequence based on the embedding module to obtain reference sequence features;
[0091] S3032: Perform position feature embedding on the sequence position of each keyword in the reference sequence based on the embedding module to obtain an important position sequence feature;
[0092] S3033: Fusing the reference sequence features and the importance position sequence features to obtain reference mapping features.
[0093] Specifically, the embedding module sets a trainable fourth parameter matrix and a fifth parameter matrix. First, the reference sequence is input into the embedding module, and each keyword is mapped to a reference word segmentation encoding ID (reference token ID). Based on the fourth parameter matrix, each reference word segmentation encoding ID is feature embedded to obtain a reference word segmentation feature (reference word embedding). Then, the reference sequence feature is obtained after splicing according to the sorting of the reference sequence. The reference sequence feature is a vectorized representation of the reference sequence. In addition, each keyword has a corresponding importance position feature (importance position embedding) in the reference sequence to retain the relative position relationship of the extracted reference sequence. The importance position sequence feature is obtained after splicing the importance position features according to the sorting of the reference sequence.
[0094] Specifically, fusion here refers to adding the reference sequence features and the importance position sequence features to generate the reference mapping features. By embedding word features and embedding features of sequence positions that carry keyword importance information, we can generate feature vectors that directly capture importance. This increases the proportion of keyword importance information in the final text representation, which is beneficial for the accuracy and specificity of text representation.
[0095] S305: Input the text mapping features and the reference mapping features into the feature encoding module of the target text model for feature encoding to obtain a long text representation result.
[0096] Specifically, the feature encoding module performs feature cross-processing on the text mapping features and the reference mapping features to achieve feature fusion extraction and obtain a long text representation result. The representation result integrates the semantic information related to the text order expressed in the text mapping features, the keyword importance information, and the importance ranking information expressed in the reference mapping features, to achieve multidimensional modeling of text information, which is beneficial to the recall accuracy and recall rate in subsequent retrieval and other task applications.
[0097] Accordingly, refer toFigure 8 In one embodiment, the long text processing method can specifically include: S21, obtaining an original long text; S22, performing text cleaning and word segmentation processing on the original long text to obtain a target long text; S23, inputting the target long text into a keyword extraction module to extract keywords to obtain a plurality of keywords and a word weight of each keyword, so as to generate word weight information; S24, sorting and splicing the plurality of keywords based on the text order of the target long text to obtain a keyword sequence; S25, sorting and splicing the plurality of keywords based on the word weight information to obtain a reference sequence; S26, performing word feature embedding on each keyword in the keyword sequence based on an embedding module of a target text model to obtain a word sequence feature; S26, performing position feature embedding on the sequence position of each keyword in the keyword sequence based on the embedding module to obtain a position sequence feature; S27, performing weight feature embedding on each word weight in the word weight information based on the embedding module to obtain a weight sequence feature; S28, adding the word sequence feature, the position sequence feature, and the weight sequence feature to obtain a text mapping feature; S29, performing word feature embedding on each keyword in the reference sequence based on the embedding module to obtain a reference sequence feature; S30, performing position feature embedding on the sequence position of each keyword in the reference sequence based on the embedding module to obtain an importance position sequence feature; S31, fusing the reference sequence feature and the importance position sequence feature to obtain a reference mapping feature; S32, inputting the text mapping feature and the reference mapping feature into a feature encoding module to extract features, and performing pooling on the feature vectors of the hidden states output by the hidden layer of the feature encoding module to obtain a long text representation result.
[0098] In specific embodiments, the reference Figure 9 The feature encoding module includes a first feature extraction layer and a second feature extraction layer, and S305 includes:
[0099] S3051: inputting the text mapping feature into the first feature extraction layer to perform feature extraction based on an attention mechanism to obtain an intermediate text feature;
[0100] S3052: inputting the intermediate text feature and the reference mapping feature into the second feature extraction layer to perform feature fusion extraction based on a cross-attention mechanism to obtain a long text representation result, and in the process of feature fusion extraction based on the cross-attention mechanism, the reference mapping feature is taken as a query feature and the intermediate text feature is taken as a key-value feature.
[0101] Specifically, the first feature extraction layer and the second feature extraction layer are constructed based on the attention mechanism, the feature cross of the attention mechanism is performed on the text mapping feature first, then the obtained intermediate text feature is taken as the key value, the reference mapping feature is taken as the query feature, and further cross-attention feature cross fusion is performed, which can organically combine the information expressed by the two source feature sequences, and accurately model the text.
[0102] In some embodiments, the first feature extraction layer is constructed based on a multi-head self-attention mechanism, and the second feature extraction layer is constructed based on a cross-attention mechanism. The first feature extraction layer is used to learn the association between feature vectors at different positions in the input feature sequence through cross processing, and the second feature extraction layer is used to learn the association between the feature vectors between the input text mapping features and the reference mapping features. The reference sequence features are sorted according to importance, so the importance position sequence features carry the importance information of the keywords. The importance information carried by the reference mapping features is further fused with the text mapping features through the cross-attention mechanism to improve the feature modeling effect.
[0103] In some embodiments, reference Figure 10 , S305 may include:
[0104] S3053: Inputting the text mapping features and the reference mapping features into the first feature extraction layer to perform fusion feature extraction based on the attention mechanism to obtain fused text features;
[0105] S3054: The fused text features and the reference mapping features are input into the second feature extraction layer for feature fusion extraction based on the cross-attention mechanism to obtain the long text representation result. In the process of feature fusion extraction based on the cross-attention mechanism, the reference mapping features are used as query features and the fused text features are used as key features.
[0106] Specifically, the text mapping features and the reference mapping features are spliced and input into the first feature extraction layer for a fusion to obtain the fused text features, which are then input into the second feature extraction layer with the reference mapping features for a second fusion extraction to obtain the long text representation results. This further increases the proportion of important keyword information carried by the reference mapping features in the representation results, thereby improving the accuracy and specificity of text feature modeling.
[0107] Based on some or all of the above embodiments, in the embodiments of this application, reference Figure 11 The method also includes a model training method, including S401-S409:
[0108] S401: Obtain an initial text model and training data. The initial text model is a dual-tower model, including two embedding modules with shared parameters and two feature encoding modules with shared parameters. The training data includes multiple sample text pairs.
[0109] S403: For each sample text pair, obtain the sample keyword sequence and sample word weight information of the two sample texts respectively, the sample keyword sequence includes a plurality of sample keywords corresponding to the sample text and is ordered based on the text order of the sample text, and the sample word weight information includes a sample word weight of each sample keyword in the plurality of sample keywords, and the sample word weight is used to represent the importance of the sample keyword to the sample text.
[0110] S405: Input the sample keyword sequence and sample word weight information of the plurality of sample text pairs into the initial text model to obtain first sample mapping features and second sample mapping features based on feature mapping of the sample keyword sequence and sample word weight information of the sample text pair by the two embedding modules respectively; and obtain the first sample text representation corresponding to the first sample mapping feature and the second sample text representation corresponding to the second sample mapping feature based on feature extraction of the first sample mapping feature and the second sample mapping feature by the two feature encoding modules respectively.
[0111] S407: Loss calculation is performed based on the similarity between the first sample text representation and the second sample text representation to obtain a model loss.
[0112] S409: The initial text model is trained based on the model loss to adjust the network parameters of the embedding module and the feature encoding module until the training end condition is met to obtain a target text model.
[0113] Specifically, referring to Figure 12 , in the dual tower structure of the initial text model, one side is used to input one sample text in the sample text pair, and the other side is used to input the other sample text in the sample text pair to output the respective sample text representation, and then loss calculation is realized. The plurality of sample text pairs of the training data can all be long text pairs, or can include part of the short text pairs, which can be selected based on the actual training effect. The sample keyword sequence and the sample word weight information are similar to the keyword sequence and the word weight information described above, the first sample mapping feature and the second sample mapping feature are similar to the text mapping feature described above, and the first sample text representation and the second sample text representation are similar to the long text representation result described above, which will not be repeated here. The sample text pair can include a positive sample text pair and a negative sample text pair. The algorithm of vector similarity can include but is not limited to cosine distance, Hamming distance, etc.
[0114] In some embodiments, an unsupervised learning method can be used to train the initial text model. Accordingly, a simCSE (Simple Contrastive Learning of Sentence Embeddings) method or the like can be used, combined with a vector similarity algorithm and a contrastive learning method to optimize training and update model parameters until the loss of this iteration is less than a preset loss, or the difference in loss between two adjacent times is less than a preset loss difference, or a preset number of iterations is reached, to determine that the training end condition is met, and the updated initial text model that meets the training end condition is determined as the target text model.
[0115] In other embodiments, a supervised learning method can be used to train the initial text model. Accordingly, the training data also includes sample labels for sample text pairs. The sample labels are used to represent the true value of the similarity between the two sample texts in the sample text pair. For example, in a binary classification, if the two texts are similar, the sample label is 1; if they are not similar, the sample label is 0. During the training process, supervised training data is used to complete the text matching task.
[0116] The training process of supervised and unsupervised methods is the same as Figure 12 As shown, sample text pairs are preprocessed and then fed into the keyword extraction module to output their respective sample keyword sequences and sample word weight information. These are then fed into the embedding module to obtain the first and second sample mapping features. These are then fed into the feature encoding module, where the output hidden layer states are pooled to obtain the first and second sample text representations. During training, the cosine distance can be used to measure the similarity between the first and second sample text representations, and the mean square error loss (MSEloss) function is used to optimize model parameters.
[0117] In some other embodiments, after S403, the model training method further includes S501: for each sample text pair, weighting multiple sample keywords of the sample text based on sample word weight information to obtain sample reference sequences of the two sample texts;
[0118] Accordingly, refer to Figure 13 S405 may include: inputting the sample reference sequences of the two sample texts into the initial text model, performing feature mapping on the two sample reference sequences of the sample text pair based on the two embedding modules, obtaining a first sample reference mapping feature and a second sample reference mapping feature, and performing feature extraction on the first sample mapping feature, the first sample reference mapping feature, the second sample mapping feature, and the second sample reference mapping feature based on the two feature encoding modules, respectively, to obtain a first sample text representation and a second sample text representation. Then, S407 and S409 are executed to obtain the target text model.
[0119] It can be understood that the sample reference mapping feature is used to represent the sample word feature and the sample importance position feature of the sample keyword in the sample reference sequence. The sample reference sequence is similar to the aforementioned reference sequence, the first sample reference mapping feature and the second sample reference mapping feature are similar to the aforementioned reference mapping feature, and the fusion and extraction process of the first sample mapping feature and the first sample reference mapping feature by each feature encoding module or the fusion and extraction process of the second sample mapping feature and the second sample reference mapping feature are consistent with the fusion and extraction process of the text mapping feature and the reference mapping feature by the aforementioned feature encoding module, and details are not repeated here.
[0120] The long text processing method in the embodiments of the present application can be applied to retrieval, text matching and the like, and the generated long text representation result can be stored in a retrieval library as a feature to be recalled, or the target text model can be directly applied to text similarity matching between two texts. For example, after preprocessing the original long text to be queried, the keyword extraction module is input to obtain the keyword sequence and the word weight information, and then the target text model is input to obtain the long text representation result, and the long text representation result matched therewith is retrieved in the vector library to realize information recall or information sorting.
[0121] The embodiments of the present application also provide a long text processing device 800, as shown in Figure 14 The structure schematic diagram of the long text processing device provided by the embodiments of the present application is shown, and the device can include the following modules: Figure 14 The structure schematic diagram of the long text processing device provided by the embodiments of the present application is shown, and the device can include the following modules:
[0122] The acquisition module 10 is used to acquire the keyword sequence and the word weight information of the target long text, the keyword sequence includes a plurality of keywords corresponding to the target long text and is ordered based on the text order of the target long text, and the word weight information includes the word weight of each keyword in the plurality of keywords, and the word weight is used to represent the importance of the keyword to the target long text;
[0123] The feature mapping module 20 is used to perform feature mapping on the keyword sequence and the word weight information to obtain the text mapping feature, and the text mapping feature is used to represent the word feature, the weight feature and the position feature of the keyword in the keyword sequence;
[0124] The feature extraction module 30 is used to perform feature extraction on the target long text based on the text mapping feature to obtain the long text representation result.
[0125] In some embodiments, the acquisition module 10 includes:
[0126] The word extraction submodule is used to extract the keywords from the target long text to obtain a plurality of keywords;
[0127] Splicing submodule: used to sort and splice multiple keywords based on the text order of the target long text to obtain a keyword sequence;
[0128] Word weight submodule: used to determine the word weight of each keyword based on the keyword's category distinction information and the frequency of occurrence of the keyword in the target long text, and obtain word weight information. The category distinction information is used to indicate the text category distinction ability of the keyword.
[0129] In some embodiments, the feature mapping module 20 includes:
[0130] Word embedding submodule: It is used to embed word features of each keyword in the keyword sequence based on the embedding module of the target text model to obtain word sequence features;
[0131] Position embedding submodule: used to embed the position feature of the sequence position of each keyword in the keyword sequence based on the embedding module to obtain the position sequence feature;
[0132] Weight embedding submodule: used to embed the weight feature of each word in the word weight information based on the embedding module to obtain the weight sequence feature;
[0133] Sequence fusion submodule: used to fuse word sequence features, position sequence features and weight sequence features to obtain text mapping features.
[0134] In some embodiments, the apparatus further comprises:
[0135] Original text acquisition module: used to obtain the original long text before obtaining the keyword sequence and word weight information of the target long text;
[0136] Preprocessing module: used to perform text cleaning and word segmentation on the original long text to obtain the target long text.
[0137] In some embodiments, the feature extraction module 30 is specifically used to: input the text mapping features into the feature encoding module of the target text model to perform feature encoding based on the attention mechanism to obtain a long text representation result.
[0138] In some other embodiments, the feature extraction module 30 includes:
[0139] Weight sorting submodule: used to sort the multiple keywords of the target long text based on word weight information to obtain a reference sequence;
[0140] Reference mapping submodule: It is used to perform feature mapping on the reference sequence based on the embedding module of the target text model to obtain reference mapping features. The reference mapping features are used to represent the word features and importance position features of the keywords in the reference sequence.
[0141] Feature encoding submodule: It is used to input text mapping features and reference mapping features into the feature encoding module of the target text model for feature encoding to obtain long text representation results.
[0142] In some embodiments, the reference mapping submodule includes:
[0143] Word embedding unit: used to embed word features for each keyword in the reference sequence based on the embedding module to obtain reference sequence features;
[0144] Importance unit: used to embed the position feature of the sequence position of each keyword in the reference sequence based on the embedding module to obtain the importance position sequence feature;
[0145] Reference fusion unit: used to fuse reference sequence features and importance position sequence features to obtain reference mapping features.
[0146] In some embodiments, the feature encoding module includes a first feature extraction layer and a second feature extraction layer, and the feature encoding submodule includes:
[0147] Feature extraction unit: used to input text mapping features into the first feature extraction layer for feature extraction based on the attention mechanism to obtain intermediate text features;
[0148] Fusion extraction unit: used to input the intermediate text features and reference mapping features into the second feature extraction layer for feature fusion extraction based on the cross-attention mechanism to obtain the long text representation result. In the process of feature fusion extraction based on the cross-attention mechanism, the reference mapping features are used as query features and the intermediate text features are used as key features.
[0149] In some embodiments, the apparatus further includes a model training module, specifically configured to:
[0150] Obtain an initial text model and training data. The initial text model is a dual-tower model, including two embedding modules with shared parameters and two feature encoding modules with shared parameters. The training data includes multiple sample text pairs.
[0151] For each sample text pair, obtain sample keyword sequences and sample word weight information for each of the two sample texts. The sample keyword sequence includes multiple sample keywords corresponding to the sample text and is sorted based on the text order of the sample text. The sample word weight information includes the sample word weight of each sample keyword in the multiple sample keywords. The sample word weight is used to represent the importance of the sample keyword to the sample text.
[0152] Inputting sample keyword sequences and sample word weight information of a plurality of sample text pairs into an initial text model, performing feature mapping on the sample keyword sequences and sample word weight information of the sample text pairs based on two embedding modules, respectively, to obtain a first sample mapping feature and a second sample mapping feature; and performing feature extraction on the first sample mapping feature and the second sample mapping feature based on two feature encoding modules, respectively, to obtain a first sample text representation corresponding to the first sample mapping feature and a second sample text representation corresponding to the second sample mapping feature;
[0153] Calculating the loss based on the similarity between the first sample text representation and the second sample text representation to obtain the model loss; and
[0154] The initial text model is trained based on the model loss to adjust the network parameters of the embedding module and the feature encoding module until the training end conditions are met to obtain the target text model.
[0155] It should be noted that the above device embodiments and method embodiments are based on the same implementation method.
[0156] An embodiment of the present application provides a device, which may be a terminal or a server, including a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the long text processing method or neural network training method provided in the above-mentioned method embodiment.
[0157] The memory can be used to store software programs and modules. The processor executes various functional applications and anomaly detection by running the software programs and modules stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, application programs required for functions, etc.; the data storage area can store data created based on the use of the device, etc. In addition, the memory can include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory can also include a memory controller to provide the processor with access to the memory.
[0158] The method embodiments provided in the embodiments of the present application can be executed in electronic devices such as mobile terminals, computer terminals, servers or similar computing devices. Figure 15 This is a hardware structure diagram of an electronic device for a long text processing method provided in an embodiment of the present application. Figure 15As shown, the electronic device 900 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 910 (the processor 910 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 930 for storing data, and one or more storage media 920 (such as one or more mass storage devices) for storing application programs 923 or data 922. Among them, the memory 930 and the storage medium 920 can be temporary storage or permanent storage. The program stored in the storage medium 920 may include one or more modules, each module may include a series of instruction operations on the electronic device. Furthermore, the central processing unit 910 can be configured to communicate with the storage medium 920 and execute a series of instruction operations in the storage medium 920 on the electronic device 900. The electronic device 900 may also include one or more power supplies 960, one or more wired or wireless network interfaces 950, one or more input and output interfaces 940, and / or one or more operating systems 921, such as Windows Server TM , Mac OS X TM , Unix TM , LinuxTM, FreeBSDTM, etc.
[0159] The input / output interface 940 can be used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by a communications provider of the electronic device 900. In one embodiment, the input / output interface 940 includes a network adapter (NIC), which can be connected to other network devices via a base station to communicate with the Internet. In one embodiment, the input / output interface 940 can be a radio frequency (RF) module for wirelessly communicating with the Internet.
[0160] It can be understood by those skilled in the art that Figure 15 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 15 More or fewer components than shown, or with Figure 15 Different configurations shown.
[0161] An embodiment of the present application also provides a computer-readable storage medium, which can be set in an electronic device to store at least one instruction or at least one program related to implementing an anomaly detection method in a method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the anomaly detection method provided by the above-mentioned method embodiment.
[0162] Optionally, in this embodiment, the storage medium may be located in at least one of a plurality of network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0163] According to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described above.
[0164] The long text processing method, device, equipment, storage medium, server, terminal and program product provided by the above-mentioned present application first obtains the keyword sequence and word weight information of the target long text, performs feature mapping on the keyword sequence and word weight information, obtains text mapping features, and then extracts features of the target long text based on the text mapping features to obtain the long text representation result. Among them, the keyword sequence includes multiple keywords corresponding to the target long text and is sorted based on the text order of the target long text. The word weight information includes the word weight of each keyword in the multiple keywords, which can represent the importance of the keyword to the target long text. The corresponding text mapping features are used to represent the word features, weight features and position features of the keywords in the keyword sequence. Through the extraction of keywords and text order sorting, the global core content is grasped, the integrity, balance and continuity of the expression of the key content of the long text are ensured, the text length limit is broken away, and the text content modeling effect and efficiency are improved; at the same time, word weight information is introduced to express the importance of keywords, effectively improving the expression accuracy and specific discrimination of the text content under the premise of realizing efficient representation extraction of long texts, and significantly improving the effective content recall rate and accuracy in text retrieval, matching and other scenarios.
[0165] It should be noted that the order of the embodiments of the present application described above is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0166] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device, equipment, and storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0167] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by instructing the relevant hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk, or an optical disk, etc.
[0168] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A long text processing method, characterized in that: The method comprises: Obtaining a keyword sequence and word weight information of a target long text, wherein the keyword sequence includes multiple keywords corresponding to the target long text and is sorted based on the text order of the target long text, and the word weight information includes a word weight of each keyword in the multiple keywords, and the word weight is used to represent the importance of the keyword to the target long text; Performing feature mapping on the keyword sequence and the word weight information to obtain text mapping features, wherein the text mapping features are used to characterize word features, weight features, and position features of the keywords in the keyword sequence; The feature extraction of the target long text is performed based on the text mapping feature to obtain a long text representation result.
2. The method according to claim 1, characterized in that The acquisition of the keyword sequence and word weight information of the target long text includes: Perform keyword extraction on the target long text to obtain a plurality of keywords; Sort and concatenate the multiple keywords based on the text order of the target long text to obtain the keyword sequence; The word weight of each keyword is determined based on the category differentiation information of the keyword and the occurrence frequency of the keyword in the target long text to obtain the word weight information. The category differentiation information is used to indicate the text category differentiation ability of the keyword.
3. The method according to claim 1, characterized in that The feature mapping of the keyword sequence and the word weight information to obtain text mapping features includes: An embedding module based on a target text model performs word feature embedding on each keyword in the keyword sequence to obtain a word sequence feature; Based on the embedding module, position feature embedding is performed on the sequence position of each keyword in the keyword sequence to obtain a position sequence feature; Performing weight feature embedding on each word weight in the word weight information based on the embedding module to obtain a weight sequence feature; The word sequence feature, the position sequence feature and the weight sequence feature are fused to obtain the text mapping feature.
4. The method according to claim 1, wherein Before obtaining the keyword sequence and word weight information of the target long text, the method further includes: Get the original long text; The original long text is cleaned and segmented to obtain the target long text.
5. The method according to any one of claims 1 to 4, characterized in that Extracting the features of the target long text based on the text mapping features to obtain a long text representation result includes: The text mapping features are input into the feature encoding module of the target text model for feature encoding based on the attention mechanism to obtain the long text representation result.
6. The method according to any one of claims 1 to 4, characterized in that Extracting the features of the target long text based on the text mapping features to obtain a long text representation result includes: sorting the keywords of the target long text by weight based on the word weight information to obtain a reference sequence; Performing feature mapping on the reference sequence based on an embedding module of a target text model to obtain reference mapping features, wherein the reference mapping features are used to characterize word features and importance position features of keywords in the reference sequence; The text mapping features and the reference mapping features are input into the feature encoding module of the target text model for feature encoding to obtain the long text representation result.
7. The method according to claim 6, characterized in that The performing feature mapping on the reference sequence to obtain reference mapping features includes: Performing word feature embedding on each keyword in the reference sequence based on the embedding module to obtain a reference sequence feature; Based on the embedding module, position feature embedding is performed on the sequence position of each keyword in the reference sequence to obtain an important position sequence feature; The reference sequence feature and the importance position sequence feature are fused to obtain the reference mapping feature.
8. The method according to claim 6, characterized in that The feature encoding module includes a first feature extraction layer and a second feature extraction layer. The feature encoding is performed using the text mapping feature and the reference mapping feature as inputs of the target text model to obtain the long text representation result. Inputting the text mapping features into the first feature extraction layer to perform feature extraction based on the attention mechanism to obtain intermediate text features; The intermediate text features and the reference mapping features are input into the second feature extraction layer for feature fusion extraction based on the cross-attention mechanism to obtain the long text representation result. In the process of feature fusion extraction based on the cross-attention mechanism, the reference mapping features are used as query features and the intermediate text features are used as key features.
9. The method according to any one of claims 1 to 4, characterized in that The method also includes a model training method, including: Obtaining an initial text model and training data, wherein the initial text model is a dual-tower model including two embedding modules with shared parameters and two feature encoding modules with shared parameters, and the training data includes multiple sample text pairs; For each sample text pair, obtaining a sample keyword sequence and sample word weight information of each of the two sample texts, wherein the sample keyword sequence includes multiple sample keywords corresponding to the sample text and is sorted based on the text order of the sample text; the sample word weight information includes a sample word weight of each sample keyword in the multiple sample keywords, and the sample word weight is used to represent the importance of the sample keyword to the sample text; Inputting the sample keyword sequences and sample word weight information of the plurality of sample text pairs into the initial text model, performing feature mapping on the sample keyword sequences and sample word weight information of the sample text pairs based on the two embedding modules, respectively, to obtain a first sample mapping feature and a second sample mapping feature; and performing feature extraction on the first sample mapping feature and the second sample mapping feature based on the two feature encoding modules, respectively, to obtain a first sample text representation corresponding to the first sample mapping feature and a second sample text representation corresponding to the second sample mapping feature; Calculating the loss based on the similarity between the first sample text representation and the second sample text representation to obtain a model loss; The initial text model is trained based on the model loss to adjust the network parameters of the embedding module and the feature encoding module until the training end condition is met to obtain the target text model.
10. A long text processing device, characterized in that: The device comprises: An acquisition module is configured to acquire a keyword sequence and word weight information of a target long text, wherein the keyword sequence includes a plurality of keywords corresponding to the target long text and is sorted based on the text order of the target long text; the word weight information includes a word weight of each keyword in the plurality of keywords, and the word weight is used to represent the importance of the keyword to the target long text; Feature mapping module: used to perform feature mapping on the keyword sequence and the word weight information to obtain text mapping features, wherein the text mapping features are used to characterize the word features, weight features and position features of the keywords in the keyword sequence; Feature extraction module: used to extract features of the target long text based on the text mapping features to obtain a long text representation result.
11. A computer-readable storage medium, characterized in that The storage medium stores at least one instruction or at least one program segment, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the long text processing method according to any one of claims 1 to 9.
12. A computer device, characterized in that: The device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program segment, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the long text processing method according to any one of claims 1 to 9.
13. A computer program product, characterized in that The computer program product includes computer instructions, and when the computer instructions are executed by a processor, the long text processing method according to any one of claims 1 to 9 is implemented.