Data processing method and device, electronic equipment and storage medium
By performing word segmentation and combination processing on Chinese corpus information, the target Chinese vocabulary list is generated, which solves the problem of crossing boundaries when Chinese processing in the prior art, and improves the rationality and semantic clarity of the Chinese vocabulary list.
Patent Information
- Application Number
- CN202311553084.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-05-23
AI Technical Summary
The existing word segmenter modules are prone to crossing boundaries when processing Chinese, resulting in the lack of true meaning of the processing results and cannot effectively adapt to the needs of Chinese scenarios.
By performing word segmentation processing on Chinese corpus information, the frequency of occurrence of word segmentation units is determined, and the Chinese word segmentation in the target word segmentation unit is combined to generate a target Chinese word list.
This method can effectively solve the problem of crossing boundaries in Chinese processing and improve the rationality and semantic clarity of Chinese vocabulary list.
Smart Images

Figure CN120031033A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning technology, and in particular to a data processing method, device, electronic device and storage medium. Background Art
[0002] The word segmenter module in the data preprocessing stage of the large model can convert the input natural language into a word list that the large model can process. The existing word segmenter module can specifically convert the input string into a string byte list and perform segmentation processing in units of bytes; since a Chinese character occupies multiple bytes, after segmentation processing in units of bytes, boundary crossing may occur. For example, when high-frequency words are merged, the third byte of the Chinese character is combined with the byte corresponding to the adjacent English character behind it. The character substring formed by this combination actually has no real meaning; that is, the existing word segmenter module cannot adapt well to the processing requirements of Chinese scenarios. Summary of the invention
[0003] The technical problem to be solved by the present application is to provide a data processing method, device, electronic device and storage medium, which can solve the cross-boundary problem that occurs when the existing word segmentation module processes Chinese, can adapt well to the processing requirements of Chinese scenarios, and improve the rationality of the Chinese vocabulary.
[0004] In order to solve the above technical problems, on the one hand, the present application provides a data processing method, including:
[0005] When the data to be processed includes Chinese corpus information, performing word segmentation processing on the Chinese corpus information to obtain a plurality of Chinese word segments;
[0006] Determine the occurrence frequencies corresponding to the plurality of word segmentation units respectively; each word segmentation unit includes a plurality of consecutive Chinese word segmentations in the Chinese corpus information; the occurrence frequency corresponding to each word segmentation unit is the number of times each word segmentation unit appears in the Chinese corpus information;
[0007] Determine a word segmentation unit whose occurrence frequency meets a first preset frequency condition among the multiple word segmentation units as a target word segmentation unit;
[0008] Combining the Chinese word segments in the target word segmentation unit to obtain a target subword;
[0009] A target Chinese word list corresponding to the Chinese corpus information is generated based on the multiple Chinese word segmentations and the target subwords.
[0010] On the other hand, the present application provides a data processing device, comprising:
[0011] A first word segmentation module is used to perform word segmentation processing on the Chinese corpus information to obtain a plurality of Chinese word segments when the data to be processed includes Chinese corpus information;
[0012] A frequency determination module, used to determine the occurrence frequencies corresponding to a plurality of word segmentation units; each word segmentation unit includes a plurality of consecutive Chinese word segments in the Chinese corpus information; the occurrence frequency corresponding to each word segmentation unit is the number of times each word segmentation unit appears in the Chinese corpus information;
[0013] A word segmentation unit determination module, configured to determine a word segmentation unit whose occurrence frequency satisfies a first preset frequency condition among the multiple word segmentation units as a target word segmentation unit;
[0014] A word segmentation combination module, used for combining the Chinese word segments in the target word segmentation unit to obtain a target subword;
[0015] A Chinese vocabulary generation module is used to generate a target Chinese vocabulary corresponding to the Chinese corpus information based on the multiple Chinese word segments and the target subwords.
[0016] On the other hand, the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the data processing method as described above.
[0017] On the other hand, the present application provides a computer storage medium, wherein the storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded by a processor and executes the data processing method as described above.
[0018] Implementing the embodiments of the present application has the following beneficial effects:
[0019] In the present application, Chinese corpus information is segmented to obtain multiple Chinese word segments, where the word segmentation is based on Chinese semantics, and the multiple Chinese word segments obtained all have clear semantics; thus, the word segmentation unit is determined based on the multiple Chinese word segments, and the target subwords obtained are combined based on the Chinese word segments in the target word segmentation unit, and the multiple continuous Chinese word segments with clear semantics will be merged, which is different from the problem in the prior art that after word segmentation by bytes, word merging results in the formation of merged words without real meaning due to the cross-boundary phenomenon. Therefore, the method of first segmenting the Chinese corpus information and then combining the Chinese word segments to generate a Chinese word list can well adapt to the processing requirements of Chinese scenarios and improve the rationality of the Chinese word list. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 It is a schematic diagram of the implementation environment provided by the embodiment of the present application;
[0022] Figure 2 is a flow chart of a data processing method provided in an embodiment of the present application;
[0023] Figure 3 is a flow chart of a method for adding words to a Chinese vocabulary provided in an embodiment of the present application;
[0024] Figure 4 is a flow chart of a method for constructing a vocabulary based on a business processing target provided in an embodiment of the present application;
[0025] Figure 5 is a flow chart of a method for updating a vocabulary based on each batch of data to be processed provided by an embodiment of the present application;
[0026] Figure 6 It is a flow chart of a method for generating a field Chinese vocabulary provided in an embodiment of the present application;
[0027] Figure 7 It is a flow chart of a data deduplication method provided in an embodiment of the present application;
[0028] Figure 8 is a schematic diagram of a data processing device provided in an embodiment of the present application;
[0029] Fig. 9 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] To make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0032] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0033] See also Figure 1 , which shows a schematic diagram of an implementation environment provided by an embodiment of the present application, the implementation environment may include: at least one user terminal 110 and a data processing terminal 120, and the user terminal 110 and the data processing terminal 120 can communicate data through a network.
[0034] Specifically, the user terminal 110 may send a data processing request to the data processing terminal 120, and the data processing request may carry data to be processed, and the data to be processed may include Chinese corpus information; the data processing request may be used to request the data processing terminal 120 to generate a Chinese vocabulary based on the Chinese corpus information. When the data processing terminal 120 receives the data processing request, it performs word segmentation processing on the Chinese corpus information in the data processing request to obtain multiple Chinese word segments; then, based on the multiple Chinese word segments, it determines the word segmentation unit, determines the frequency of occurrence of the word segmentation unit, and combines the Chinese word segments in the target word segmentation unit, thereby generating a target Chinese vocabulary; and further, the generated target Chinese vocabulary may be returned to the user terminal 110.
[0035] The user terminal 110 can communicate with the data processing end 120 based on the browser / server mode (B / S) or the client / server mode (C / S). The user terminal 110 may include: physical devices such as smart phones, tablet computers, laptops, digital assistants, smart wearable devices, and vehicle terminals, and may also include software running in the physical devices, such as applications. The operating system running on the user terminal 110 in the embodiment of the present application may include but is not limited to Android system, IOS system, Linux, Windows, etc.
[0036] The data processing end 120 can establish a communication connection with the user terminal 110 via wired or wireless communication. The data processing end 120 can include an independently operated server, or a distributed server, or a server cluster composed of multiple servers, wherein the server can be a cloud server.
[0037] The large language model can be a pre-training model, also known as a cornerstone model or a large model. It refers to a deep neural network (DNN) with large parameters. It is trained on a large amount of unlabeled data. The function approximation ability of the large parameter DNN is used to enable the PTM to extract common features from the data. After fine-tuning, parameter efficient fine-tuning (PEFT), prompt-tuning and other technologies, it is suitable for downstream tasks. Therefore, the pre-training model can achieve ideal results in a few-shot or zero-shot scenario. According to the data modality processed, PTM can be divided into language models (ELMO, BERT, GPT), visual models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multimodal models (ViBERT, CLIP, Flamingo, Gato), etc., among which the multimodal model refers to a model that establishes two or more data modality feature representations. The pre-training model is an important tool for outputting artificial intelligence generated content (AIGC), and can also be used as a general interface to connect multiple specific task models.
[0038] In order to solve the problem that the word segmentation module in the prior art crosses the boundary when processing Chinese and cannot adapt well to the processing requirements of the Chinese scene, the embodiment of the present application provides a data processing method, and the execution subject of the method can be the above-mentioned data processing end; please refer to Figure 2 , the method may specifically include:
[0039] S210. When the data to be processed includes Chinese corpus information, perform word segmentation processing on the Chinese corpus information to obtain a plurality of Chinese word segments.
[0040] The data to be processed in this embodiment may be the original training data obtained during the training of the large language model. Before the large language model is trained based on the training data, the original training data may be preprocessed. The data preprocessing may include data deduplication, text segmentation, and other operations. For the preprocessing operation of data deduplication, it is possible to perform deduplication processing on the original training data, thereby reducing the amount of data that needs to be processed by the large language model during the training process, and at the same time, it is possible to avoid the large language model from learning similar training data and avoid wasting training resources; for the preprocessing operation of text segmentation, the original training data is difficult to be recognized by the large language model, so that after the preprocessing operation of text segmentation is performed on the original training data, the word list obtained through text segmentation can be directly recognized and processed by the large language model.
[0041] The data to be processed may include Chinese corpus information, Western corpus information and other multi-language corpus information; for Western corpus information, such as English, since each English letter occupies one byte, the English word segmentation can be directly performed based on bytes. For Chinese, each Chinese character can occupy multiple bytes. If it is segmented in bytes, there will be a phenomenon of crossing boundaries when high-frequency words are merged, and the original semantics of Chinese will be lost. Therefore, in this embodiment, the Chinese corpus information can be segmented to obtain multiple Chinese word segments; specifically, when performing word segmentation, word segmentation based on character granularity can be used, or word segmentation based on word granularity can be used; here, the word segmentation of the Chinese corpus information can be implemented using existing Chinese word segmentation tools, such as jieba, HanLP, FoolNLTK and other Chinese word segmentation tools, which are not specifically limited in this embodiment.
[0042] S220. Determine the occurrence frequencies corresponding to a plurality of word segmentation units; each word segmentation unit includes a plurality of consecutive Chinese word segments in the Chinese corpus information; the occurrence frequency corresponding to each word segmentation unit is the number of times each word segmentation unit appears in the Chinese corpus information.
[0043] The word segmentation unit can be determined based on the order of multiple Chinese word segmentations in the Chinese corpus information, and then the frequency of occurrence of each word segmentation unit can be determined, and each word segmentation unit includes multiple continuous Chinese word segmentations in the Chinese corpus information. For example, the Chinese corpus information includes "You like to travel, and he also likes to travel". The Chinese corpus information is segmented according to the word granularity to obtain "You", "Like", "Travel", "He", "Also", "Like", "Travel". If each word segmentation unit includes two continuous Chinese word segmentations in the Chinese corpus information, multiple word segmentation units "You, like", "Like, Travel", "Travel, He", "He, Also", "Also, Like", "Like, Travel" can be obtained accordingly. It can be seen that the number of times the word segmentation units "You, like", "Travel, He", "He, Also", "Also, Like" appear in the Chinese corpus information is 1, and the frequency of occurrence of these word segmentation units is 1. The number of times the word segmentation unit "Like, Travel" appears in the Chinese corpus information is 2, and the frequency of occurrence of this word segmentation unit is 2. Furthermore, a word segmentation unit including three consecutive Chinese word segmentations, or a word segmentation unit including four consecutive Chinese word segmentations, etc. can also be determined. In this embodiment, the occurrence frequency of word segmentation units including different numbers of Chinese word segmentations can be counted, so as to determine the target word segmentation unit for Chinese word segmentation combination based on the occurrence frequency of various types of word segmentation units.
[0044] S230. Determine a word segmentation unit whose occurrence frequency meets a first preset frequency condition among the multiple word segmentation units as a target word segmentation unit.
[0045] The first preset frequency condition in this embodiment may be the highest frequency of occurrence or the frequency of occurrence being greater than or equal to the first preset frequency; thus, the target word segmentation unit may be the word segmentation unit with the highest frequency of occurrence among multiple word segmentation units, such as the above-mentioned word segmentation unit "like, travel"; the target word segmentation unit may be the word segmentation unit with the frequency of occurrence being greater than or equal to the first preset frequency among multiple word segmentation units, and if the first preset frequency is 2, then the target word segmentation unit may be the above-mentioned word segmentation unit "like, travel"; if the first preset frequency is 1, then the target word segmentation unit may include the above-mentioned "you, like", "like, travel", "travel, him", "he, also", "also, like". The target word segmentation unit may be one or more.
[0046] In a specific embodiment, there may be multiple word segmentation units with equal frequency of occurrence. If the multiple word segmentation units with equal frequency of occurrence all meet the first preset frequency condition, the multiple word segmentation units with equal frequency of occurrence can be determined as target word segmentation units. It should be noted that the multiple word segmentation units with equal frequency of occurrence may include multiple word segmentation units that do not have a containment relationship, such as word segmentation unit "A, B", word segmentation unit "C, D"; they may also include multiple word segmentation units that have a containment relationship, such as word segmentation unit "E, F", word segmentation unit "E, F, G", in which case the word segmentation unit "E, F" is included in the word segmentation unit "E, F, G", and in the absence of special instructions, multiple word segmentation units that have a containment relationship can be used as target word segmentation units; wherein A, B, C, D, E, F, G each corresponds to a Chinese word segmentation.
[0047] S240. Combining the Chinese word segments in the target word segmentation units to obtain target subwords.
[0048] When combining multiple Chinese word segments in each target word segmentation unit, corresponding target subwords can be generated according to the arrangement order of the multiple Chinese word segmentations in the Chinese corpus information; for example, for the above-mentioned word segmentation unit "like, travel", in the Chinese corpus information, "like" is located before "travel", so the generated target subword is "like to travel".
[0049] S250. Generate a target Chinese vocabulary corresponding to the Chinese corpus information based on the multiple Chinese word segmentations and the target subwords.
[0050] Multiple Chinese word segments can be used as basic words in the Chinese word list. If there are repeated word segments in the multiple Chinese word segments, the multiple Chinese word segments can be deduplicated to obtain deduplicated Chinese word segments, and the deduplicated Chinese word segments can be used as basic words in the Chinese word list. On the basis of determining the basic words, target subwords can also be added to the Chinese word list to generate a target Chinese word list corresponding to the Chinese corpus information.
[0051] It should be noted that, when the determined target Chinese vocabulary is larger than the upper limit of the number of words in the Chinese vocabulary, the words in the target Chinese vocabulary can be deleted; the deletion operation can be determined based on the frequency of occurrence of each word in the target Chinese vocabulary in the Chinese corpus information. Specifically, the words with a frequency of occurrence less than a preset frequency can be deleted, or the words with a preset number of words ranked lower can be determined by sorting from high to low based on the frequency of occurrence. The deletion operation can also be implemented based on words with a containment relationship. For multiple words with a containment relationship, if the frequency of occurrence of these multiple words in the Chinese corpus information is the same, the included words can be deleted. For example, the above-mentioned word segmentation unit "E, F" and word segmentation unit "E, F, G" generate corresponding subwords EF and EFG, and the subwords EF and EFG have the same frequency of occurrence in the Chinese corpus information. At this time, the subword EF can be deleted.
[0052] In the present application, Chinese corpus information is segmented to obtain multiple Chinese word segments, where the word segmentation is based on Chinese semantics, and the multiple Chinese word segments obtained all have clear semantics; thus, the word segmentation unit is determined based on the multiple Chinese word segments, and the target subwords obtained are combined based on the Chinese word segments in the target word segmentation unit, and the multiple continuous Chinese word segments with clear semantics will be merged, which is different from the problem in the prior art that after word segmentation by bytes, word merging results in the formation of merged words without real meaning due to the cross-boundary phenomenon. Therefore, the method of first segmenting the Chinese corpus information and then combining the Chinese word segments to generate a Chinese word list can well adapt to the processing requirements of Chinese scenarios and improve the rationality of the Chinese word list.
[0053] After learning the Chinese corpus information, if the number of words in the current Chinese vocabulary has not reached the upper limit of the number of words in the Chinese vocabulary, that is, the target number, you can perform the word addition operation on the Chinese vocabulary; please refer to Figure 3 , which shows a method for adding words to a Chinese vocabulary, the method may include:
[0054] S310. When the number of words in the Chinese vocabulary determined based on the multiple Chinese word segmentations and the target subwords is less than the target number, determine the Chinese vocabulary determined based on the multiple Chinese word segmentations and the target subwords as the current Chinese vocabulary.
[0055] In the case where the multiple Chinese word segmentations do not include repeated word segmentations, the Chinese word list can be directly determined based on the multiple Chinese word segmentations and the target subword; in the case where the multiple Chinese word segmentations include repeated word segmentations, the multiple Chinese word segmentations can be first deduplicated, and then the Chinese word list can be determined based on the deduplicated Chinese word segmentations and the target subword. In the case where the number of words in the determined Chinese word list is less than the target number, it means that words need to be added to the Chinese word list.
[0056] S320. Determine a current word segmentation unit; the current word segmentation unit is a word segmentation unit other than the target word segmentation unit among the multiple word segmentation units; the occurrence frequency of the current word segmentation unit meets a second preset frequency condition.
[0057] The second preset frequency condition in this embodiment may be the highest occurrence frequency or the occurrence frequency greater than or equal to the second preset frequency; thus, the current word segmentation unit may be the word segmentation unit with the highest occurrence frequency among multiple word segmentation units except the target word segmentation unit; the current word segmentation unit may also be a word segmentation unit with an occurrence frequency greater than or equal to the second preset frequency; the current word segmentation unit may be one or more.
[0058] In a specific embodiment, there may be multiple word segmentation units with equal frequency of occurrence. If the multiple word segmentation units with equal frequency of occurrence all meet the second preset frequency condition, the multiple word segmentation units with equal frequency of occurrence can be determined as the current word segmentation unit. It should be noted that the multiple word segmentation units with equal frequency of occurrence may include multiple word segmentation units that do not have a containment relationship, such as word segmentation unit "A, B", word segmentation unit "C, D"; they may also include multiple word segmentation units that have a containment relationship, such as word segmentation unit "E, F", word segmentation unit "E, F, G", in which case the word segmentation unit "E, F" is included in the word segmentation unit "E, F, G", and in the absence of special instructions, multiple word segmentation units that have a containment relationship can be used as the current word segmentation unit; wherein A, B, C, D, E, F, G each corresponds to a Chinese word segmentation.
[0059] S330. Combining the Chinese word segments in the current word segmentation unit to obtain the current subword.
[0060] When combining multiple Chinese word segments in each current word segmentation unit, corresponding target subwords can be generated according to the arrangement order of the multiple Chinese word segmentations in the Chinese corpus information; for example, for the above-mentioned word segmentation unit "you, like", in the Chinese corpus information, "you" is located before "like", so the generated target subword is "you like".
[0061] S340. Add the current subword to the current Chinese word list, and determine the current word segmentation unit as the target word segmentation unit.
[0062] The newly generated current subword is added to the current Chinese vocabulary to update the current Chinese vocabulary, and the current word segmentation unit is determined as the target word segmentation unit, because the target word segmentation unit will not be repeatedly determined as the current word segmentation unit, thereby avoiding repeated determination of the current word segmentation unit when determining the current word segmentation unit in the next cycle.
[0063] S350. Determine whether the number of Chinese words in the current Chinese word list is greater than or equal to the target number; if so, execute step S360; if not, execute step S320.
[0064] If the number of Chinese words in the current Chinese vocabulary is greater than or equal to the target number, it means that the current Chinese vocabulary has reached the upper limit of the number of words, and you can stop adding words to the Chinese vocabulary; if the number of Chinese words in the current Chinese vocabulary is less than the target number, it means that the current Chinese vocabulary has not reached the upper limit of the number of words, and you can continue to add new words to the Chinese vocabulary.
[0065] S360. Determine the current Chinese vocabulary as the target Chinese vocabulary.
[0066] Therefore, after learning the Chinese corpus information, if the number of words in the Chinese vocabulary has not reached the upper limit of the number of words, high-frequency word segmentation units can be further determined from the word segmentation units that have not been determined as target word segmentation units in each cycle, and the Chinese word segmentations in the high-frequency word segmentation units determined in the current cycle are combined to obtain new sub-words; this method can achieve the expansion of the number of words in the Chinese vocabulary without increasing the Chinese corpus, thereby improving the convenience of expanding the Chinese vocabulary.
[0067] Large language models usually need to be applicable to multi-language processing scenarios, so it is necessary to generate vocabulary tables corresponding to multiple languages. According to different business processing requirements, the requirements for vocabulary tables in different languages are different, so vocabulary tables can be constructed based on business processing goals; please refer to Figure 4 , which shows a vocabulary construction method based on business processing objectives, which may include:
[0068] S410. Based on the business processing objectives, determine the vocabulary sizes corresponding to the multiple language types; the vocabulary size corresponding to each language type is the upper limit of the number of words in the vocabulary of each language type.
[0069] The vocabulary size corresponding to each language type can be set in advance as a hyperparameter, and can be flexibly adjusted based on actual business needs. By adjusting the vocabulary size of each language type, the understanding and modeling capabilities of the large language model for different language types can be adjusted. For example, by increasing the vocabulary size of the Chinese vocabulary, the modeling ability of the large language model for Chinese can be improved. The language types in this embodiment may include Chinese, Western and other language types. Specifically, the total vocabulary size of multiple language types is 128,000, of which the vocabulary size of the English vocabulary can be set to 32,000. For the Chinese vocabulary, the number of Chinese characters is limited to 20,933 commonly used characters, and the corresponding number of Chinese words can be 75,067.
[0070] S420. When each batch of data to be processed is obtained, determine the target language type of each batch of data to be processed.
[0071] Each batch of data to be processed may include corpus information, and the target language type of the data to be processed can be determined based on the specific content of the corpus information.
[0072] S430. When the number of words in the current vocabulary corresponding to the target language type is less than the upper limit of the number of words of the target language type, update the current vocabulary corresponding to the target language type based on the to-be-processed data of the target language type.
[0073] Specifically, the size of the vocabulary can be limited by adding a conditional judgment during the training process. In each round of inputting new corpus, it is determined whether the vocabulary of the target language type has reached the vocabulary size based on the target language type of the new corpus; if the corresponding vocabulary size has not been reached, the new corpus of the target language type can be received and processed to obtain new words, and the new words are added to the current vocabulary to achieve vocabulary update.
[0074] In a specific embodiment, when the number of words in the current vocabulary corresponding to the target language type is greater than or equal to the upper limit of the number of words of the target language type, the to-be-processed data of the target language type is ignored, that is, the to-be-processed data of the target language type is no longer processed; and a feedback prompt message can also be issued to remind the operating user not to continue to provide the to-be-processed data of the target language type, and the vocabulary of the target language type has reached a preset vocabulary size.
[0075] In this embodiment, through the business processing objectives, the vocabulary size corresponding to each language can be flexibly set. For the language type that needs to be processed emphatically, the vocabulary size corresponding to the language type can be appropriately increased to improve the natural language processing and modeling capabilities of the large language model for the language type; further, by processing the data to be processed in batches, before processing each batch of the data to be processed, it is first determined based on the language type whether data processing is required. If so, the processing continues; if not, the data to be processed may not be processed, and a prompt is given that there is no need to provide or obtain the data to be processed of the target language type in the future. This avoids the problem of constructing a vocabulary based on the full amount of data to be processed at one time, while the corresponding vocabulary has been constructed before some of the data to be processed in the full amount of data to be processed has not been processed. This realizes the on-demand acquisition of data to be processed of each language type, and avoids the waste of resources caused by the acquired data to be processed not being processed.
[0076] It should be noted that, if the data to be processed in the current batch includes corpus information of multiple language types at the same time, if the vocabularies of multiple language types have not reached the vocabulary scale, the vocabularies of multiple language types can be updated simultaneously based on the data to be processed in the current batch; if the vocabulary of at least one language type among the multiple language types reaches the vocabulary scale, the corpus information of at least one language type is ignored, and the vocabularies of other language types are updated based on other corpus information in the current batch, wherein the other language types are language types other than at least one language type among the multiple language types in the current batch, and the other corpus information is the corpus information corresponding to the other language types in the current batch.
[0077] Furthermore, the vocabulary can be updated based on each batch of data to be processed; see Figure 5 , which shows a method for updating a vocabulary based on each batch of data to be processed, including:
[0078] S510. When the current batch of data to be processed includes Chinese corpus information and the number of words in the current Chinese vocabulary is less than the upper limit of the number of words in the Chinese vocabulary, the Chinese corpus information of the current batch is segmented to obtain multiple Chinese word segments of the current batch.
[0079] Specifically, when performing word segmentation, word segmentation based on character granularity or word segmentation based on word granularity can be used; the word segmentation processing of Chinese corpus information here can be implemented using existing Chinese word segmentation tools, such as jieba, HanLP, FoolNLTK and other Chinese word segmentation tools, which are not specifically limited in this embodiment.
[0080] S520. Determine the occurrence frequencies corresponding to the multiple word segmentation units of the current batch; each word segmentation unit includes multiple consecutive Chinese word segments in the Chinese corpus information of the current batch; the occurrence frequency corresponding to each word segmentation unit is the number of times each word segmentation unit appears in the Chinese corpus information of the current batch.
[0081] The word segmentation unit can be determined based on the order of multiple Chinese word segmentations in the Chinese corpus information, and then the frequency of occurrence of each word segmentation unit can be determined, and each word segmentation unit includes multiple continuous Chinese word segmentations in the Chinese corpus information. For example, the Chinese corpus information includes "You like to travel, and he also likes to travel". The Chinese corpus information is segmented according to the word granularity to obtain "You", "Like", "Travel", "He", "Also", "Like", "Travel". If each word segmentation unit includes two continuous Chinese word segmentations in the Chinese corpus information, multiple word segmentation units "You, like", "Like, Travel", "Travel, He", "He, Also", "Also, Like", "Like, Travel" can be obtained accordingly. It can be seen that the number of times the word segmentation units "You, like", "Travel, He", "He, Also", "Also, Like" appear in the Chinese corpus information is 1, and the frequency of occurrence of these word segmentation units is 1. The number of times the word segmentation unit "Like, Travel" appears in the Chinese corpus information is 2, and the frequency of occurrence of this word segmentation unit is 2. Furthermore, a word segmentation unit including three consecutive Chinese word segmentations, or a word segmentation unit including four consecutive Chinese word segmentations, etc. can also be determined. In this embodiment, the occurrence frequency of word segmentation units including different numbers of Chinese word segmentations can be counted, so as to determine the word segmentation unit for Chinese word segmentation combination based on the occurrence frequency of various types of word segmentation units.
[0082] S530. Determine the adjacent word segments whose occurrence frequencies satisfy the third preset frequency condition among the multiple Chinese word segments of the current batch as the target word segment units of the current batch.
[0083] The third preset frequency condition in this embodiment may be the highest frequency of occurrence or the frequency of occurrence being greater than or equal to the third preset frequency; thus, the target word segmentation unit of the current batch may be the word segmentation unit with the highest frequency of occurrence among the multiple word segmentation units of the current batch, such as the above-mentioned word segmentation unit "like, travel"; the target word segmentation unit of the current batch may be the word segmentation unit with a frequency of occurrence greater than or equal to the third preset frequency among the multiple word segmentation units of the current batch, if the third preset frequency is 2, then the target word segmentation unit of the current batch may be the above-mentioned word segmentation unit "like, travel"; if the third preset frequency is 1, then the target word segmentation unit of the current batch may include the above-mentioned "you, like", "like, travel", "travel, him", "he, also", "also, like". The target word segmentation unit of the current batch may also be the word segmentation unit with a frequency of occurrence greater than or equal to the third preset frequency, for example, if the third preset frequency is 2, the corresponding target word segmentation unit is "like, travel"; the target word segmentation unit of the current batch may be one or more.
[0084] In a specific embodiment, there may be multiple word segmentation units with equal frequency of occurrence. If the multiple word segmentation units with equal frequency of occurrence all meet the third preset frequency condition, the multiple word segmentation units with equal frequency of occurrence can be determined as the target word segmentation units of the current batch. It should be noted that the multiple word segmentation units with equal frequency of occurrence may include multiple word segmentation units that do not have a containment relationship, such as word segmentation unit "A, B", word segmentation unit "C, D"; they may also include multiple word segmentation units that have a containment relationship, such as word segmentation unit "E, F", word segmentation unit "E, F, G", in which case the word segmentation unit "E, F" is included in the word segmentation unit "E, F, G", and in the absence of special instructions, multiple word segmentation units that have a containment relationship can be used as the target word segmentation units of the current batch; wherein A, B, C, D, E, F, G each corresponds to a Chinese word segmentation.
[0085] S540. Combining the Chinese word segments in the word segmentation units corresponding to the current batch to obtain the subwords of the current batch.
[0086] When combining multiple Chinese word segments in each target word segmentation unit of the current batch, sub-words of the current batch can be generated according to the arrangement order of the multiple Chinese word segmentations in the Chinese corpus information; for example, for the above-mentioned word segmentation unit "like, travel", in the Chinese corpus information, "like" is located before "travel", so the sub-word generated for the current batch is "like to travel".
[0087] S550. Update the current Chinese word list based on the multiple Chinese word segments of the current batch and the subwords of the current batch to obtain an updated Chinese word list.
[0088] The multiple Chinese word segments of the current batch can be used as basic words in the Chinese word list of the current batch. Furthermore, if there are repeated word segments in the multiple Chinese word segments of the current batch, the multiple Chinese word segments of the current batch can be deduplicated to obtain deduplicated Chinese word segments, and the deduplicated Chinese word segments can be used as basic words in the Chinese word list. On the basis of determining the basic words, the subwords of the current batch can also be added to the Chinese word list to generate the Chinese word list of the current batch.
[0089] In this embodiment, the data to be processed can be processed in batches. In the process of processing each batch of data to be processed, Chinese word segmentation, word segmentation unit determination, Chinese word segmentation merging of high-frequency word segmentation units and other operations can be performed based on the data to be processed in this batch to generate corresponding new words, and the current Chinese vocabulary list can be updated based on the new words. The amount of data in each batch of data to be processed is relatively small compared to the total amount of data to be processed, so the amount of data required to process operations such as Chinese word segmentation, word segmentation unit determination, Chinese word segmentation merging of high-frequency word segmentation units and other operations is also small, thereby improving the efficiency of generating new words, and correspondingly improving the efficiency of updating the Chinese vocabulary list.
[0090] Furthermore, the Chinese vocabulary may include Chinese characters and Chinese words. For the Chinese characters, this embodiment provides a method for generating a Chinese vocabulary based on the Chinese characters. The method may include:
[0091] Acquire a plurality of preset Chinese characters; the plurality of preset Chinese characters include common Chinese characters and uncommon Chinese characters;
[0092] The Chinese vocabulary is generated based on the common Chinese characters, the uncommon Chinese characters, the Chinese word segmentations, and the target subwords.
[0093] If Chinese characters are obtained only through Chinese corpus information, there may be some rare Chinese characters that are not covered by the corpus, which results in rare Chinese characters not being included in the Chinese vocabulary, and then the large language model cannot process the rare Chinese characters; in this embodiment, commonly used Chinese characters are written into the Chinese vocabulary, and on this basis, the Chinese word segmentation and target subwords obtained after processing the Chinese corpus information can be combined to generate a Chinese vocabulary. In this way, rare Chinese characters can be included in the Chinese vocabulary, ensuring that the large language model has the ability to process rare Chinese characters.
[0094] When it is necessary to generate a corresponding Chinese vocabulary for a specific target domain, operations such as Chinese word segmentation, word segmentation unit determination, and Chinese word segmentation merging of high-frequency word segmentation units can be performed based on the Chinese corpus information of the target domain. For details, please refer to Figure 6 , which shows a method for generating a domain Chinese vocabulary, including:
[0095] S610. Based on the multiple Chinese word segmentations, the target subwords and the general Chinese vocabulary, a target Chinese vocabulary corresponding to the target domain is obtained by merging.
[0096] The Chinese corpus information in this embodiment may specifically be the Chinese corpus information of the target domain, so that based on the implementation method of steps S210 to S250 of this embodiment, a plurality of Chinese word segments and target subwords corresponding to the target domain can be obtained. The general Chinese vocabulary may be a general Chinese vocabulary generated after performing operations such as Chinese word segmentation, word segmentation unit determination, and Chinese word segmentation merging of high-frequency word segmentation units based on the general Chinese corpus information; the general Chinese vocabulary may also be obtained through the implementation method of steps S210 to S250 of this embodiment.
[0097] The Chinese word segmentations, target subwords and general Chinese word lists corresponding to the target domain are merged to obtain the corresponding target Chinese word list. During the word list merging process, the Chinese word segmentations and target subwords corresponding to the target domain can be retained, and low-frequency Chinese words in the general Chinese word list can be deleted.
[0098] S620. Obtain incremental Chinese corpus information in the target domain.
[0099] The incremental Chinese corpus information is different from the Chinese corpus information in the target field, and the data volume of the incremental Chinese corpus information is smaller than the data volume of the Chinese corpus information in the target field, that is, a small amount of Chinese corpus information in the target field is additionally obtained.
[0100] S630. Update the target Chinese vocabulary corresponding to the target domain based on the incremental Chinese corpus information to obtain an updated target Chinese vocabulary corresponding to the target domain.
[0101] Based on a small amount of Chinese corpus information in the target domain, training operations are performed to optimize the newly added words.
[0102] Thus, by processing the Chinese corpus information based on the target domain, words for the target domain are obtained, and combined with the general Chinese vocabulary adapted to the general domain, the target Chinese vocabulary corresponding to the target domain is generated, which realizes the general vocabulary capability of the large language model itself, and enables the large language model to be applicable to specific target domain scenarios, improving the multi-domain adaptation capability of the large language model. In addition, after the words corresponding to the target domain are merged with the general Chinese vocabulary, additional training based on a small amount of Chinese corpus information can further improve the adaptability of the Chinese vocabulary to the target domain.
[0103] The preprocessing operation for data deduplication can achieve deduplication of the original training data, thereby reducing the amount of data that needs to be processed by the large language model during the training process, and at the same time avoiding the large language model from learning similar training data, thereby avoiding the waste of training resources; before deduplication of the Chinese corpus, it is also necessary to segment the Chinese corpus, and perform feature extraction based on the segmentation results, calculate the similarity of the Chinese corpus based on the extracted features, and then perform data deduplication based on the similarity. Since a Chinese character occupies multiple bytes, after being segmented in units of bytes, each segmented word lacks actual semantics. Therefore, after feature extraction based on the segmentation word, the extracted features are difficult to represent the original Chinese semantic information, which leads to inaccurate similarity calculation, seriously affecting the recall rate and effect of deduplication. In order to solve this problem, this embodiment provides a data deduplication method, please refer to Figure 7 , the method may include:
[0104] S710. Extract local features of the text based on the multiple Chinese word segmentations to obtain a first local feature corresponding to the Chinese corpus information in the data to be processed.
[0105] When performing word segmentation in this embodiment, word segmentation based on character granularity or word granularity can be adopted; the word segmentation processing of Chinese corpus information here can be implemented by using existing Chinese word segmentation tools, such as jieba, HanLP, FoolNLTK and other Chinese word segmentation tools, which are not specifically limited in this embodiment.
[0106] N-Gram is a text feature representation method that can be used to capture local context information in a text so as to compare the similarity between texts. In this embodiment, 2-Gram or 3-Gram can be used. Taking 2-Gram as an example, the Chinese corpus information is "I bought a computer today". The multiple Chinese word segments obtained by word segmentation at word granularity include "I", "today", "bought", "one", and "computer". 2-Gram can be represented as a set ["I today", "bought today", "bought one", "one computer"], and then the first local feature of the Chinese corpus information can be determined based on 2-Gram.
[0107] S720. Determine the feature similarity between a first local feature corresponding to the Chinese corpus information in the data to be processed and a second local feature corresponding to the historical Chinese corpus information; the second local feature is obtained based on text local feature extraction of multiple Chinese word segmentations corresponding to the historical Chinese corpus information.
[0108] Similarly, the second local feature corresponding to the historical Chinese corpus information can also be implemented using the above-mentioned N-Gram, and the first local feature and the second local feature use the same N-Gram, for example, 2-Gram or 3-Gram are used at the same time.
[0109] In one example, when calculating the feature similarity between the first local feature and the second local feature, the set similarity between the N-Gram set 1 corresponding to the Chinese corpus information in the data to be processed and the N-Gram set 2 corresponding to the historical corpus information can be specifically calculated. In this embodiment, the ratio between the intersection and the union of the two sets of N-Gram set 1 and N-Gram set 2 can be calculated, and the calculated ratio is between 0 and 1, and the closer the ratio is to 1, the more similar the two sets are.
[0110] In another example, when comparing the similarity of two sets, the set elements in N-Gram set 1 and N-Gram set 2 can be hashed using a random permutation function, and the minimum value output by the hash function is selected as the signature of the set, that is, the similarity of the two sets is converted into the similarity of the two set signatures, and the similarity of the two sets is determined by judging the similarity of the two set signatures, thereby achieving efficient set similarity comparison.
[0111] S730. When the feature similarity is greater than or equal to a preset similarity, deduplication processing is performed on the Chinese corpus information in the data to be processed.
[0112] In this embodiment, the preset similarity can be 0.85 or 0.9, etc., which can be determined according to the specific implementation situation; if the feature similarity is greater than or equal to the preset similarity, the Chinese corpus information in the data to be processed is deleted, that is, the Chinese corpus information in the data to be processed will not be used in the process of constructing the Chinese vocabulary; when the feature similarity is less than the preset similarity, the Chinese corpus information in the data to be processed is retained to construct the Chinese vocabulary based on the Chinese corpus information in the data to be processed.
[0113] When word segmentation is performed at the word granularity, the deduplication algorithm is more sensitive to Chinese repetitions, and can avoid the problem of the model's ability to recall repeated data being reduced due to changes in some words in long sentences.
[0114] In the data deduplication process, this embodiment first performs word segmentation on the Chinese corpus information to obtain multiple Chinese word segments with specific semantics, and then performs text local feature extraction based on the Chinese word segmentation, so that the extracted local features carry the original Chinese semantics, thereby improving the accuracy of local feature representation, and then performs feature similarity calculation based on the accurate local features to determine whether it is necessary to perform deduplication processing on the Chinese corpus information in the data to be processed, thereby improving the accuracy and robustness of data deduplication.
[0115] The present application provides a semantic-based Chinese word segmentation process for Chinese corpus information during data deduplication and vocabulary building for Chinese corpus data, which solves the problem that the existing technology of segmenting Chinese by bytes is not suitable for Chinese scenarios, so that the large model can better adapt to the processing requirements of Chinese scenarios.
[0116] See also Figure 8 This embodiment further provides a data processing device, including:
[0117] A first word segmentation module 810 is used to perform word segmentation processing on the Chinese corpus information to obtain a plurality of Chinese word segments when the data to be processed includes Chinese corpus information;
[0118] The frequency determination module 820 is used to determine the occurrence frequencies corresponding to the plurality of word segmentation units; each word segmentation unit includes a plurality of consecutive Chinese word segments in the Chinese corpus information; the occurrence frequency corresponding to each word segmentation unit is the number of times each word segmentation unit appears in the Chinese corpus information;
[0119] A word segmentation unit determination module 830, configured to determine a word segmentation unit whose occurrence frequency satisfies a first preset frequency condition among the multiple word segmentation units as a target word segmentation unit;
[0120] A word segmentation combination module 840, configured to combine Chinese word segments in the target word segmentation units to obtain target subwords;
[0121] The Chinese vocabulary generation module 850 is used to generate a target Chinese vocabulary corresponding to the Chinese corpus information based on the multiple Chinese word segmentations and the target subwords.
[0122] Furthermore, the Chinese vocabulary generation module 850 includes:
[0123] A first determining module, configured to determine the Chinese vocabulary determined based on the multiple Chinese word segmentations and the target subwords as the current Chinese vocabulary when the number of words in the Chinese vocabulary determined based on the multiple Chinese word segmentations and the target subwords is less than the target number;
[0124] A second determination module is used to determine a current word segmentation unit; the current word segmentation unit is a word segmentation unit other than the target word segmentation unit among the multiple word segmentation units; the occurrence frequency of the current word segmentation unit meets a second preset frequency condition;
[0125] A first combining module, configured to combine the Chinese words in the current word segmentation unit to obtain a current subword;
[0126] An adding module, used for adding the current subword to the current Chinese word list, and determining the current word segmentation unit as the target word segmentation unit;
[0127] A repeating module is used to repeatedly execute the steps of: determining a current word segmentation unit; combining Chinese words in the current word segmentation unit to obtain a current subword; adding the current subword to the current Chinese word list and determining the current word segmentation unit as the target word segmentation unit; until the number of Chinese words in the current Chinese word list is greater than or equal to the target number;
[0128] The third determination module is used to determine the current Chinese vocabulary at the end of the loop as the target Chinese vocabulary.
[0129] Furthermore, the device also includes:
[0130] A vocabulary size determination module, used to determine the vocabulary size corresponding to each of the multiple language types based on the business processing target; the vocabulary size corresponding to each language type is the upper limit of the number of words in the vocabulary of each language type;
[0131] A target language type determination module, used to determine the target language type of each batch of data to be processed when each batch of data to be processed is obtained;
[0132] The first updating module is configured to update the current vocabulary corresponding to the target language type based on the to-be-processed data of the target language type when the number of words in the current vocabulary corresponding to the target language type is less than the upper limit of the number of words of the target language type.
[0133] Furthermore, the first update module includes:
[0134] A second word segmentation module is used to perform word segmentation processing on the Chinese corpus information of the current batch to obtain a plurality of Chinese word segments of the current batch when the data to be processed in the current batch includes Chinese corpus information and the number of words in the current Chinese vocabulary is less than the upper limit of the number of words in the Chinese vocabulary;
[0135] A fourth determination module is used to determine the occurrence frequencies corresponding to the multiple word segmentation units of the current batch; each word segmentation unit includes multiple consecutive Chinese word segments in the Chinese corpus information of the current batch; the occurrence frequency corresponding to each word segmentation unit is the number of times each word segmentation unit appears in the Chinese corpus information of the current batch;
[0136] A fifth determination module, configured to determine the word segmentation units whose occurrence frequency satisfies the third preset frequency condition among the multiple word segmentation units of the current batch as target word segmentation units of the current batch;
[0137] A second combining module, configured to combine the Chinese words in the target word segmentation units of the current batch to obtain the subwords of the current batch;
[0138] The second updating module is used to update the current Chinese word list based on the multiple Chinese word segments of the current batch and the subwords of the current batch to obtain an updated Chinese word list.
[0139] Furthermore, the device also includes:
[0140] A preset Chinese character acquisition module, used to acquire a plurality of preset Chinese characters; the plurality of preset Chinese characters include common Chinese characters and uncommon Chinese characters;
[0141] The Chinese vocabulary generation module includes:
[0142] The first generation module is used to generate the Chinese vocabulary based on the common Chinese characters, the rare Chinese characters, the Chinese word segmentations, and the target subwords.
[0143] Furthermore, the Chinese corpus information is Chinese corpus information of the target domain;
[0144] The Chinese vocabulary generation module includes:
[0145] A merging module, configured to merge the multiple Chinese word segments, the target subwords and the general Chinese word list to obtain a target Chinese word list corresponding to the target domain;
[0146] An incremental information acquisition module, used to acquire incremental Chinese corpus information of the target domain;
[0147] The second generating module is used to update the target Chinese vocabulary corresponding to the target domain based on the incremental Chinese corpus information to obtain an updated target Chinese vocabulary corresponding to the target domain.
[0148] Furthermore, the device also includes:
[0149] A feature extraction module, configured to extract local features of the text based on the multiple Chinese word segmentations, and obtain a first local feature corresponding to the Chinese corpus information in the data to be processed;
[0150] A feature similarity determination module, used to determine the feature similarity between a first local feature corresponding to the Chinese corpus information in the data to be processed and a second local feature corresponding to the historical Chinese corpus information; the second local feature is obtained by extracting local features of text from a plurality of Chinese word segments corresponding to the historical Chinese corpus information;
[0151] The deduplication module is used to perform deduplication processing on the Chinese corpus information in the data to be processed when the feature similarity is greater than or equal to a preset similarity.
[0152] The device provided in the above embodiment can execute the method provided in any embodiment of the present application, and has the corresponding functional modules and beneficial effects of executing the method. For technical details not described in detail in the above embodiment, please refer to the method provided in any embodiment of the present application.
[0153] This embodiment further provides a computer-readable storage medium, in which at least one instruction or at least one program is stored. The at least one instruction or at least one program is loaded by a processor and executed as any of the above methods of this embodiment.
[0154] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs any of the above methods.
[0155] Fig. 9 is a block diagram of an electronic device for data processing according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as shown in FIG. Fig. 9 As shown. The electronic device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a data processing method is implemented.
[0156] Those skilled in the art will understand that Fig. 9The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present disclosure, and does not constitute a limitation on the electronic device to which the scheme of the present disclosure is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0157] This specification provides method operation steps as described in the embodiments or flow charts, but more or fewer operation steps may be included based on conventional or non-creative labor. The steps and sequence listed in the embodiments are only one way of executing the sequence of many steps and do not represent the only execution sequence. When the actual system or interrupt product is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, in a parallel processor or multi-threaded processing environment).
[0158] The structure shown in this embodiment is only a partial structure related to the scheme of the present application, and does not constitute a limitation on the device to which the scheme of the present application is applied. The specific device may include more or fewer components than shown, or combine certain components, or have different arrangements of components. It should be understood that the method, device, etc. disclosed in this embodiment can be implemented in other ways. For example, the device embodiment described above is only schematic. For example, the division of the modules is only a division of a logical function. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or unit modules.
[0159] Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., various media that can store program codes.
[0160] Those skilled in the art may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0161] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A data processing method, It is characterized in that include: When the data to be processed includes Chinese corpus information, performing word segmentation processing on the Chinese corpus information to obtain a plurality of Chinese word segments; Determine the occurrence frequencies corresponding to the plurality of word segmentation units respectively; each word segmentation unit includes a plurality of consecutive Chinese word segmentations in the Chinese corpus information; the occurrence frequency corresponding to each word segmentation unit is the number of times each word segmentation unit appears in the Chinese corpus information; Determine a word segmentation unit whose occurrence frequency meets a first preset frequency condition among the multiple word segmentation units as a target word segmentation unit; Combining the Chinese word segments in the target word segmentation unit to obtain a target subword; A target Chinese word list corresponding to the Chinese corpus information is generated based on the multiple Chinese word segmentations and the target subwords.
2. The method according to claim 1, It is characterized in that The generating a target Chinese vocabulary corresponding to the Chinese corpus information based on the Chinese word segmentation and the target subwords includes: When the number of words in the Chinese vocabulary determined based on the multiple Chinese word segmentations and the target subwords is less than the target number, determining the Chinese vocabulary determined based on the multiple Chinese word segmentations and the target subwords as the current Chinese vocabulary; Determine a current word segmentation unit; the current word segmentation unit is a word segmentation unit other than the target word segmentation unit among the multiple word segmentation units; the occurrence frequency of the current word segmentation unit meets a second preset frequency condition; Combining the Chinese word segments in the current word segmentation unit to obtain a current subword; Adding the current subword to the current Chinese word list, and determining the current word segmentation unit as the target word segmentation unit; Repeat the steps of: determining a current word segmentation unit; combining Chinese words in the current word segmentation unit to obtain a current subword; adding the current subword to the current Chinese word list, and determining the current word segmentation unit as the target word segmentation unit; until the number of Chinese words in the current Chinese word list is greater than or equal to the target number; The current Chinese vocabulary at the end of the loop is determined as the target Chinese vocabulary.
3. The method according to claim 1, It is characterized in that The method further comprises: Based on the business processing objectives, determine the vocabulary size corresponding to each of the multiple language types; the vocabulary size corresponding to each language type is the upper limit of the number of words in the vocabulary of each language type; When each batch of data to be processed is obtained, determining the target language type of each batch of data to be processed; When the number of words in the current vocabulary corresponding to the target language type is less than the upper limit of the number of words of the target language type, the current vocabulary corresponding to the target language type is updated based on the to-be-processed data of the target language type.
4. The method according to claim 3, It is characterized in that When the number of words in the current vocabulary corresponding to the target language type is less than the upper limit of the number of words of the target language type, updating the current vocabulary corresponding to the target language type based on the to-be-processed data of the target language type includes: When the current batch of data to be processed includes Chinese corpus information and the number of words in the current Chinese vocabulary is less than the upper limit of the number of words in the Chinese vocabulary, performing word segmentation processing on the current batch of Chinese corpus information to obtain multiple Chinese word segments of the current batch; Determine the occurrence frequencies corresponding to the multiple word segmentation units of the current batch; each word segmentation unit includes multiple consecutive Chinese word segments in the Chinese corpus information of the current batch; the occurrence frequency corresponding to each word segmentation unit is the number of times each word segmentation unit appears in the Chinese corpus information of the current batch; Determine the word segmentation units whose occurrence frequency meets the third preset frequency condition among the multiple word segmentation units of the current batch as the target word segmentation units of the current batch; Combining the Chinese word segments in the target word segmentation units of the current batch to obtain the subwords of the current batch; The current Chinese word list is updated based on the multiple Chinese word segments of the current batch and the subwords of the current batch to obtain an updated Chinese word list.
5. The method according to claim 1, It is characterized in that The method further comprises: Acquire a plurality of preset Chinese characters; the plurality of preset Chinese characters include common Chinese characters and uncommon Chinese characters; The step of generating a Chinese word list corresponding to the Chinese corpus information based on the Chinese word segmentation and the target subword includes: The Chinese vocabulary is generated based on the common Chinese characters, the uncommon Chinese characters, the Chinese word segmentations, and the target subwords.
6. The method according to claim 1, It is characterized in that The Chinese corpus information is Chinese corpus information in the target domain; The step of generating a target Chinese vocabulary corresponding to the Chinese corpus information based on the multiple Chinese word segmentations and the target subwords includes: Merging the multiple Chinese word segments, the target subwords, and a general Chinese word list to obtain a target Chinese word list corresponding to the target domain; Acquire incremental Chinese corpus information of the target domain; The target Chinese vocabulary corresponding to the target domain is updated based on the incremental Chinese corpus information to obtain an updated target Chinese vocabulary corresponding to the target domain.
7. The method according to claim 1, It is characterized in that After the Chinese corpus information is segmented to obtain a plurality of Chinese segmented words, the method further includes: Extract local features of the text based on the multiple Chinese word segmentations to obtain a first local feature corresponding to the Chinese corpus information in the data to be processed; Determine the feature similarity between a first local feature corresponding to the Chinese corpus information in the data to be processed and a second local feature corresponding to the historical Chinese corpus information; the second local feature is obtained by extracting local features of text from a plurality of Chinese word segments corresponding to the historical Chinese corpus information; When the feature similarity is greater than or equal to a preset similarity, deduplication processing is performed on the Chinese corpus information in the data to be processed.
8. A data processing device, It is characterized in that include: A first word segmentation module is used to perform word segmentation processing on the Chinese corpus information to obtain a plurality of Chinese word segments when the data to be processed includes Chinese corpus information; A frequency determination module, used to determine the occurrence frequencies corresponding to a plurality of word segmentation units; each word segmentation unit includes a plurality of consecutive Chinese word segments in the Chinese corpus information; the occurrence frequency corresponding to each word segmentation unit is the number of times each word segmentation unit appears in the Chinese corpus information; A word segmentation unit determination module, configured to determine a word segmentation unit whose occurrence frequency satisfies a first preset frequency condition among the multiple word segmentation units as a target word segmentation unit; A word segmentation combination module, used for combining the Chinese word segments in the target word segmentation unit to obtain a target subword; A Chinese vocabulary generation module is used to generate a target Chinese vocabulary corresponding to the Chinese corpus information based on the multiple Chinese word segments and the target subwords.
9. An electronic device, It is characterized in that The device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the data processing method according to any one of claims 1 to 7.
10. A computer storage medium, It is characterized in that The storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded by a processor and executes the data processing method according to any one of claims 1 to 7.