Data processing method and device, electronic equipment and storage medium

By dynamically dividing and identifying strings and removing backstory strings, the problem of high cost of data storage devices is solved, and the storage utilization rate is improved and the enterprise operation cost is reduced.

CN120371826APending Publication Date: 2025-07-25INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510495026.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

With the development of cloud computing and big data, the high cost of data storage devices is the problem of how to improve storage utilization to reduce the number and capacity requirements of devices.

Method used

Dividing strings through dynamic division methods of different character lengths, identifying and removing duplicate palindromic strings, and optimizing the data storage structure.

Benefits of technology

It improves storage utilization, reduces the number and capacity requirements of storage devices, and reduces the operating costs of enterprises.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371826A_ABST
    Figure CN120371826A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, electronic equipment and a storage medium, and relates to the technical field of computers, character strings are divided in multiple dynamic division modes with different character lengths, palintext character strings in sub-character strings are found out, the division mode containing the most palintext character strings is used as a target division mode, and the target division mode is used for dividing the palintext character strings in the sub-character strings. According to the method, repeated contents in the data can be found to the greatest extent by maximizing the number of the palindromic character strings, so that repeated parts in the character strings can be accurately identified and removed, and the characteristics of the palindromic character strings are fully utilized to optimize a data storage structure. Therefore, the technical problem of high cost of data storage equipment caused by increase of the total data amount can be solved, and the technical effects of improving the storage utilization rate, reducing the number and capacity requirements of the storage equipment and further reducing the operation cost of enterprises are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method, apparatus, electronic device, and storage medium for data processing. Background Art

[0002] With the integration of cloud computing, big data, and artificial intelligence into all walks of life, the total amount of data has grown explosively, and the amount of data that enterprises and organizations need to store is also increasing continuously, resulting in a relatively high cost of data storage devices. How to improve storage utilization can reduce the number and capacity requirements of storage devices, thereby reducing the operating costs of enterprises, which is an urgent problem to be solved currently. Summary of the Invention

[0003] This application provides a method, apparatus, electronic device, and storage medium for data processing to at least solve the problem of relatively high costs of data storage devices in related technologies.

[0004] This application provides a method for data processing, including:

[0005] Obtaining the starting positions of different strings;

[0006] Dynamically partitioning the strings in sequence according to different dynamic partitioning methods to obtain substrings, and determining the palindromic strings in the substrings; wherein, the character lengths of different partitioning methods are different;

[0007] Determining the partitioning method containing the most palindromic strings as the target partitioning method, and performing partitioning on the string based on the target partitioning method to achieve deduplication of the string; wherein, the partitioning methods of different strings are different.

[0008] This application also provides a data processing apparatus, including:

[0009] An obtaining unit, configured to obtain the starting positions of different strings;

[0010] A dynamic partitioning unit, configured to dynamically partition the strings in sequence according to different dynamic partitioning methods to obtain substrings, and determine the palindromic strings in the substrings; wherein, the character lengths of different partitioning methods are different;

[0011] A determining unit, configured to determine the partitioning method containing the most palindromic strings as the target partitioning method, and perform partitioning on the string based on the target partitioning method to achieve deduplication of the string; wherein, the partitioning methods of different strings are different.

[0012] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above data processing apparatuses when executing the computer program.

[0013] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above data processing methods are implemented.

[0014] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the above data processing methods are implemented.

[0015] Through the present application, since the string is divided in a dynamic division manner with multiple different character lengths, the palindromic strings in the substring are found, and the division manner containing the most palindromic strings is used as the target division manner. By maximizing the number of palindromic strings, the duplicate content in the data can be discovered to the greatest extent. In this way, the duplicate parts in the string can be accurately identified and removed, and the palindromic string characteristics are fully utilized to optimize the data storage structure. Therefore, the technical problem of high cost of data storage devices caused by the growth of the total amount of data can be solved, and the technical effects of improving the storage utilization rate, reducing the number and capacity requirements of storage devices, and further reducing the enterprise operation cost can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a schematic flowchart of a data processing method provided by an embodiment of the present application;

[0018] Figure 2 It is a schematic diagram of dynamic division provided by an embodiment of the present application;

[0019] Figure 3 It is a schematic structural diagram of a data processing device provided by an embodiment of the present disclosure;

[0020] Figure 4 It is a schematic structural diagram of another data processing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0022] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0023] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific implementation manners.

[0024] An embodiment of this application provides a method for processing data. In combination with the execution process of the method for processing data, the method will be described in detail.

[0025] Figure 1 It is a schematic flowchart of a method for processing data provided by an embodiment of this application.

[0026] As Figure 1 shown, the method includes the following steps:

[0027] Step 101, obtain the starting positions of different strings.

[0028] In an actual data processing scenario, the input data often contains multiple strings. Specifically, a storage structure such as a list or an array can be preset to record the starting position information. When scanning the data, starting from the first character of the data, according to certain judgment rules, such as character encoding features, specific character combinations or format identifiers, etc., the starting point of each string is identified. For example, when the data contains multiple types of strings, each type starts with a specific prefix, then during the scanning process, once the prefix character is detected, its position will be recorded as the starting position of the corresponding string. This way can quickly and effectively locate the starting points of different strings, providing accurate basic data for subsequent processing of strings according to different dynamic partitioning methods. At the same time, this method based on feature judgment and recording can adapt to various complex data formats and string types, ensuring the flexibility and reliability of the operation of obtaining the starting positions, and ensuring that subsequent operations such as dynamic partitioning of strings, determination of palindromic strings, and deduplication can be carried out smoothly based on accurate starting positions.

[0029] Step 102, dynamically partition the strings in turn according to different dynamic partitioning methods to obtain substrings, and determine the palindromic strings in the substrings; wherein, the character lengths of different partitioning methods are different.

[0030] In some embodiments, different dynamic partitioning methods are based on different character lengths. For example, partitioning is performed according to a character length of 1, or partitioning is performed according to a character length of 3. Specifically, the embodiments of the present application do not limit this.

[0031] In the specific partitioning process, the dynamic partitioning of the string can be started from the minimum character length. Taking the minimum character length as the starting benchmark for the partitioning operation, the character length is incremented. Each time it is incremented, the dynamic partitioning of the string is performed again until all dynamic partitioning methods are completed, and the partitioning is finished. This method can analyze the string from multiple different perspectives by gradually changing the partitioning character length, and obtain a series of substrings of different lengths.

[0032] Step 103, determine the partitioning method that contains the most palindromic strings as the target partitioning method, and perform the partitioning of the string based on the target partitioning method to achieve deduplication of the string; wherein, the partitioning methods for different strings are different.

[0033] The partitioning method that contains the most palindromic strings is selected because the palindromic string itself has a repetitive and symmetric structural feature, and this partitioning method can maximize the discovery of potential repetitive patterns in the string. Based on this target partitioning method, the string is partitioned again. During the partitioning process, since the repetitive parts represented by the palindromic strings will be recognized, these repetitive contents can be removed, thereby achieving deduplication of the string. It should be clear that due to the differences in the structure, content, etc. of different strings, the partitioning methods that can discover the most palindromic strings applicable to each of them are also different. Therefore, different strings will adopt different partitioning methods that are adapted to them to ensure that the deduplication operation can achieve the best effect in various string processing scenarios.

[0034] The deduplication storage of palindromic strings is not exactly the same. If the length of the palindromic string is odd, for example, the string 1234321, then it is necessary to store the string 1234 and the length 7 of the entire string, so that it is known that the deduplicated part is the string 321; if the length of the palindromic string is even, for example, the string 12344321, then it is necessary to store the string 1234 and the length 8 of the entire string, so that it is known that the deduplicated part is the string 4321. That is, if the length of the palindromic string is m, if m is odd, the deduplicated part is the last (n - 1) / 2 - bit string; if m is even, the deduplicated part is the last n / 2 - bit string.

[0035] Through this application, by dividing the string in a dynamic division method with multiple different character lengths, finding the palindrome string in the substring, and taking the division method containing the most palindrome strings as the target division method, by maximizing the number of palindrome strings, it is possible to find the duplicate content in the data to the greatest extent, so that the duplicate parts in the string can be accurately identified and removed, and the characteristics of the palindrome string can be fully utilized to optimize the data storage structure. Therefore, the technical problem of high cost of data storage equipment caused by the increase in the total amount of data can be solved, and the technical effect of improving storage utilization, reducing the number of storage devices and capacity requirements, and thus reducing the operating costs of enterprises can be achieved.

[0036] In some embodiments, obtaining the starting positions of different character strings further includes:

[0037] The starting positions of different character strings are determined based on multiple loops; wherein at least one loop traverses the length of the character string, and at least one loop traverses the starting position of the subsequence.

[0038] In this data processing method, in order to accurately and comprehensively determine the starting positions of different character strings, a method based on multiple loops is adopted.

[0039] Multiple loops are a structured algorithmic strategy that consists of at least two layers of loops nested within each other. One of the loops is dedicated to iterating over the length of a string. The length of a string determines the bounds of the string. By iterating over the length of the string, the algorithm is able to start from the first character of the string and gradually move to the last character, thus gaining a comprehensive understanding of the overall size of the string. During the traversal process, every character position is considered to ensure that no possible starting point of the string is missed. For example, for a string of length n, the loop will start at position 0 and increment to position n-1, thus covering the range of all possible starting positions within the string.

[0040] Another loop is focused on traversing the starting position of the subsequence. After determining the overall length range of the string, it is necessary to further determine the starting point of the subsequence. By traversing the starting position of the subsequence, the algorithm can start from every possible position of the string and try to find potential subsequences. This traversal method can analyze the string from different angles and discover hidden subsequence features. For example, in a string containing multiple words, by traversing the starting position of the subsequence, the starting point of each word can be accurately identified.

[0041] These two nested loops cooperate with each other to form an orderly and comprehensive search mechanism. The outer loop controls the traversal of the string length, and the inner loop traverses the starting positions of subsequences within each possible length range. In this way, every possible starting position of a subsequence is taken into account, without missing any potential starting points. The advantage of this method lies in its systematicness and completeness, ensuring that subsequent dynamic partitioning and analysis operations on the string can be carried out based on accurate starting positions.

[0042] By determining the starting positions of different strings based on multiple loops, it provides an accurate starting point for subsequent processing of the strings, making the entire data processing flow more efficient and accurate. This method has broad application prospects in fields such as data mining, natural language processing, and information retrieval.

[0043] In some embodiments, the dynamically partitioning the string according to different dynamic partitioning methods in sequence to obtain substrings includes:

[0044] Performing dynamic partitioning of the string based on the minimum character length, and incrementing the character length by the minimum step size, and re-performing dynamic partitioning of the string until the character length reaches the character length of the string.

[0045] Specifically, the dynamic partitioning starts with an operation on the string based on the minimum character length. The minimum character length is the initial partitioning granularity preset according to data characteristics and processing requirements, providing an initial fine-grained perspective for the entire partitioning process and ensuring that any possible short sequence features in the string are not missed.

[0046] After completing the initial partitioning based on the minimum character length, it enters the dynamic adjustment stage, that is, incrementing the character length by the minimum step size. The minimum step size is also set according to the actual processing scenario. It stipulates the minimum amplitude of the change in the partitioning granularity each time, ensuring that the partitioning process can progress step by step without skipping important intermediate states due to too large a step size. After each increment of the character length, the dynamic partitioning operation on the string is re-performed with the new character length standard. This means that after each change in the partitioning granularity, the string will be segmented from a new perspective, thus obtaining a set of substrings with different length combinations.

[0047] This process of dynamic adjustment and re - division is continuously repeated until the character length reaches the character length of the string itself. When the character length is equal to the total length of the string, the entire dynamic division process reaches the coarsest granularity, completing the full - range coverage from the finest granularity to the coarsest granularity. This way of gradually increasing the character length and repeating the division is like observing the string through multi - dimensional slicing, enabling a comprehensive and systematic acquisition of substrings of the string at different division granularities. These substrings contain the structural information of the string at different levels, which not only helps to discover local features in the string but also provides a rich data basis for subsequent operations such as determining palindromic substrings in the substrings and selecting the optimal division method.

[0048] For example, when processing a text string, a smaller character - length division may capture the features of individual characters or short words. As the character length gradually increases, the divided substrings can cover longer phrases, clauses, or even complete sentence structures. Through this dynamic division method, both the detailed features of short sequences and the overall structures of long sequences can be fully explored and analyzed, thus meeting the diverse requirements of different data - processing tasks for string information and providing comprehensive and accurate basic data for subsequent data processing and analysis.

[0049] Please refer to Figure 2 , Figure 2 a schematic diagram of dynamic division provided by an embodiment of the present application. As Figure 2 shown, assume there is a string "bcdcee". Then, according to different dynamic division methods, the string is dynamically divided in sequence. When divided with a character length of 1, the substrings obtained are: b, c, d, c, e, e, where each substring is a palindromic character; when divided with a character length of 2, the substrings obtained are: bc, cd, dc, ce, ee, where "ee" is a palindromic character; when divided with a character length of 3, the substrings obtained are: bcd, cdc, dce, cee, where "cdc" is a palindromic character; when divided with a character length of 4, the substrings obtained are: bcdc, cdce, dce, where there is no palindromic character; when divided with a character length of 5, the substring obtained is: bcdcee, where there is no palindromic character. Then, it can be known that the divided strings are b, cdc, ee.

[0050] In some embodiments, determining the palindromic substrings in the substrings includes:

[0051] Determine the palindromic substring in the substring based on the character length of the substring and a preset rule; wherein, when the length of the substring is a preset value, determine the substring as a palindromic substring; when the length of the substring is greater than the preset value, determine the substring as a palindromic substring when the first and last elements of the substring are the same and the remaining subsequence is a palindrome.

[0052] When the length of the substring is equal to the preset value, it can be directly determined that the substring is a palindromic substring. This preset value is determined after analyzing the data characteristics and palindrome features. For example, some specific substrings with a length of 1 or 2, due to their simple structure and conforming to the definition of a palindrome, can be directly recognized. This direct determination method improves the processing efficiency and avoids complex judgments for simple cases.

[0053] When the length of the substring is greater than the preset value, the judgment process is relatively complex. First, it is necessary to check whether the first and last elements of the substring are the same, because a palindromic substring has a symmetric structure, and the same first and last elements are one of its important features. If the first and last elements are different, then the substring is definitely not a palindromic substring. If the first and last elements are the same, then it is necessary to further determine whether the remaining subsequence is a palindrome.

[0054] The remaining subsequence refers to the part of the substring remaining after removing the first and last elements. The remaining subsequence is judged by recursion or other methods, combined with the condition that the first and last elements are the same, and finally it is determined whether the substring is a palindromic substring. This method of judging by cases, adopting different judgment strategies according to the different lengths of the substring, not only ensures the accuracy of the judgment but also improves the processing efficiency, enabling palindromic substrings to be efficiently found when processing substrings of different lengths.

[0055] In some embodiments, after determining the palindromic substring in the substring, the method further includes:

[0056] Save the palindromic substring to a preset array.

[0057] The preset array is a data structure that has pre-allocated memory space and is used to specifically store palindromic substrings. It plays a role in data caching and management in the entire data processing process.

[0058] Saving the palindromic substring to the preset array has multiple meanings. From the perspective of subsequent processing, it provides a reference basis for further determining whether longer substrings are palindromes. When encountering a substring with a length greater than the preset value and it is necessary to determine whether its remaining subsequence is a palindrome, the remaining subsequence can be directly compared with the palindromic substrings in the preset array, avoiding repeated judgments of the already determined palindromic substrings, thus significantly improving the judgment efficiency and reducing the waste of computing resources.

[0059] From the perspective of data management, the preset array centrally stores all the determined palindromic strings, making the data more orderly and facilitating subsequent unified analysis and processing of palindromic strings. For example, it is convenient to count information such as the number and length distribution of palindromic strings, providing strong support for determining the target partitioning method that contains the most palindromic strings.

[0060] In some embodiments, when the length of the substring is greater than the preset value, and when the first and last elements of the substring are the same and the remaining subsequence is a palindromic string, determining that the substring is a palindromic string further includes:

[0061] Comparing the remaining subsequence with the palindromic strings in the preset array to determine whether the remaining subsequence is a palindromic string.

[0062] In the data processing method, when the length of the substring is greater than the preset value and the first and last elements are the same, and it is necessary to determine whether the remaining subsequence is a palindromic string, comparing the remaining subsequence with the palindromic strings in the preset array is a key operation. The preset array, as a data set storing the determined palindromic strings, plays an important role in this step. Since various previously determined palindromic strings have been saved in the array, comparing the remaining subsequence with it can utilize the existing calculation results and avoid repeated judgments.

[0063] This comparison operation traverses the preset array and matches the character sequences of the remaining subsequence with each palindromic string in the array. Once it is found that the remaining subsequence is exactly the same as a certain palindromic string in the preset array, it can be quickly determined that the remaining subsequence is a palindromic string, and then it is determined that the substring containing the remaining subsequence is a palindromic string. This method effectively reduces the computational complexity of determining palindromic strings, improves the data processing efficiency. Especially when processing a large amount of string data, by reusing the existing judgment results of palindromic strings, it significantly reduces the system resource consumption and ensures the efficient operation of the data processing process.

[0064] In some embodiments, after determining that the partitioning method that contains the most palindromic strings is the target partitioning method, and performing the partitioning of the string based on the target partitioning method to achieve deduplication of the string, the method further includes:

[0065] Storing the deduplicated string.

[0066] After determining the target partitioning method that contains the most palindromic strings and performing deduplication on the string based on this method, storing the deduplicated string is an important final step in the data processing flow. The storage operation saves the string data that has been deduplicated and had redundancy eliminated to a specified storage medium, such as a hard disk or cloud storage, according to a pre-set data storage format and storage path. This operation not only ensures the security and persistence of the processed data, preventing data loss due to unexpected situations such as system failures or power outages, but also provides a basis for subsequent data retrieval, analysis, and applications. By storing the deduplicated string in a specific location, subsequent data processing systems, analysis tools, or other applications can easily access this data and carry out further data mining, machine learning model training, data visualization, etc. based on the clean deduplicated data, giving full play to the value of the data and realizing a complete closed-loop of data processing from raw data acquisition, analysis and processing to storage and application.

[0067] The operating system manages the storage location and method of files to ensure data consistency and security; the file system allocates appropriate storage space and updates the metadata; the controller receives the write command from the operating system and converts it into instructions that the hard disk can understand, controlling the movement of the magnetic head or flash memory unit; the hard disk uses a checksum to detect errors during data writing; the file system updates the metadata of the file, recording the actual storage location of the data and other relevant information to complete the write operation.

[0068] The present invention proposes a method and device for improving storage capacity based on a dynamic programming algorithm. Its implementation logic is to judge whether the data is a palindromic string through the dynamic programming algorithm before the data is actually written to the storage device, quickly and accurately deduplicate the data, which not only improves the read and write performance of the data but also achieves the purpose of improving storage utilization. In an ideal state, the storage utilization rate of this method can be increased infinitely close to 100%. First, data preparation is carried out. The data can be text, image, audio, video, or any other form of binary information; then, according to the dynamic programming algorithm, it is judged whether the data is a palindromic string. If it meets the palindromic string, the data is processed according to the deduplication technical means of palindromic strings; finally, data writing is completed. The operating system manages the storage location and method of files to ensure data consistency and security; the file system allocates appropriate storage space and updates the metadata; the controller receives the write command from the operating system and converts it into instructions that the hard disk can understand, controlling the movement of the magnetic head or flash memory unit; the hard disk uses a checksum to detect errors during data writing; the file system updates the metadata of the file, recording the actual storage location of the data and other relevant information to complete the write operation.

[0069] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0070] The present application also provides a data processing device. Figure 3 As shown in the structural schematic diagram of a data processing device provided by an embodiment of the present disclosure, Figure 3 it includes:

[0071] An acquisition unit 21, configured to acquire the starting positions of different strings;

[0072] A dynamic partitioning unit 22, configured to sequentially perform dynamic partitioning on the string according to different dynamic partitioning methods to obtain substrings, and determine the palindromic strings in the substrings; wherein, the character lengths of different partitioning methods are different;

[0073] A determination unit 23, configured to determine the partitioning method including the most palindromic strings as the target partitioning method, and perform partitioning on the string based on the target partitioning method to achieve deduplication of the string; wherein, the partitioning methods of different strings are different.

[0074] With the present application, since the string is partitioned by multiple dynamic partitioning methods with different character lengths, the palindromic strings in the substrings are found, and the partitioning method including the most palindromic strings is used as the target partitioning method. By maximizing the number of palindromic strings, the duplicate content in the data can be discovered to the greatest extent. In this way, the duplicate parts in the string can be accurately identified and removed, and the palindromic string characteristics can be fully utilized to optimize the data storage structure. Therefore, the technical problem of high cost of data storage devices caused by the growth of the total amount of data can be solved, and the technical effects of improving the storage utilization rate, reducing the number and capacity requirements of storage devices, and further reducing the enterprise operation cost can be achieved.

[0075] Further, in a possible implementation manner of an embodiment of the present disclosure, the acquisition unit 21 is further configured to:

[0076] Determine the starting positions of different strings based on multiple loops; wherein, at least one loop traverses the string length, and at least one loop traverses the starting positions of subsequences.

[0077] Further, in a possible implementation manner of an embodiment of the present disclosure, the dynamic partitioning unit 22 is further configured to:

[0078] Perform dynamic partitioning on the string based on the minimum character length, increment the character length with the minimum step size, and re-perform dynamic partitioning on the string until the character length reaches the character length of the string.

[0079] Further, in a possible implementation manner of the embodiments of the present disclosure, the dynamic partitioning unit 22 is further configured to:

[0080] Determine a palindromic substring in the substring based on the character length of the substring and a preset rule; wherein, when the length of the substring is a preset value, determine that the substring is a palindromic substring; when the length of the substring is greater than the preset value, determine that the substring is a palindromic substring when the first and last elements of the substring are the same and the remaining subsequence is a palindrome.

[0081] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 4 shown, the apparatus further includes:

[0082] A saving unit 24, configured to save the palindromic substring to a preset array after the dynamic partitioning unit 22 determines the palindromic substring in the substring.

[0083] Further, in a possible implementation manner of the embodiments of the present disclosure, the dynamic partitioning unit 22 is further configured to:

[0084] Compare the remaining subsequence with the palindromic substrings in the preset array to determine whether the remaining subsequence is a palindromic substring.

[0085] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 4 shown, the apparatus further includes:

[0086] A storage unit 25, configured to store the deduplicated string after determining that the partitioning method including the most palindromic substrings is the target partitioning method and performing partitioning on the string based on the target partitioning method to implement deduplication of the string.

[0087] For the description of the features in the corresponding embodiments of the data processing apparatus, reference may be made to the relevant descriptions in the corresponding embodiments of the data processing method, which will not be elaborated here one by one.

[0088] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-mentioned embodiments of the data processing method.

[0089] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above-mentioned embodiments of the data processing method when running.

[0090] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media capable of storing computer programs such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs.

[0091] An embodiment of the present application also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the data processing method.

[0092] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the data processing method.

[0093] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0094] The above has introduced in detail a data processing method, device, electronic device, and storage medium provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for processing data, characterized in that, Including: Obtain the starting positions of different strings; Dynamically partition the string successively according to different dynamic partitioning methods to obtain substrings, and determine the palindromic strings in the substrings; wherein, the character lengths of different partitioning methods are different; Determine the partitioning method that contains the most palindromic strings as the target partitioning method, and perform the partitioning of the string based on the target partitioning method to achieve deduplication of the string; wherein, the partitioning methods of different strings are different.

2. The data processing method according to claim 1, wherein The obtaining of the starting positions of different strings further includes: Determine the starting positions of different strings based on multiple loops; wherein, at least one loop traverses the string length, and at least one loop traverses the starting positions of subsequences.

3. The data processing method according to claim 1, characterized in that The successively dynamically partitioning the string according to different dynamic partitioning methods to obtain substrings includes: Perform dynamic partitioning of the string based on the minimum character length, increment the character length with the minimum step size, and re-perform dynamic partitioning of the string until the character length reaches the character length of the string.

4. The data processing method according to claim 3, characterized in that The determining of the palindromic strings in the substrings includes: Determine the palindromic strings in the substrings based on the character length of the substrings and preset rules; wherein, when the length of the substring is a preset value, determine the substring as a palindromic string; when the length of the substring is greater than the preset value, when the first and last elements of the substring are the same and the remaining subsequence is a palindrome, determine the substring as a palindromic string.

5. The data processing method according to claim 4, wherein After determining the palindromic strings in the substrings, the method further includes: Save the palindromic strings to a preset array.

6. The data processing method according to claim 5, wherein The when the length of the substring is greater than the preset value, when the first and last elements of the substring are the same and the remaining subsequence is a palindromic string, determining the substring as a palindromic string further includes: Compare the remaining subsequence with the palindromic strings in the preset array to determine whether the remaining subsequence is a palindromic string.

7. The method for processing data according to any one of claims 1-6, characterized in that After determining the partitioning method that contains the most palindromic strings as the target partitioning method, and performing the partitioning of the string based on the target partitioning method to achieve deduplication of the string, the method further includes: Store the deduplicated string.

8. A data processing device, characterized in that, Including: An obtaining unit, configured to obtain the starting positions of different strings; A dynamic partitioning unit, configured to successively dynamically partition the string according to different dynamic partitioning methods to obtain substrings, and determine the palindromic strings in the substrings; wherein, the character lengths of different partitioning methods are different; A determining unit, configured to determine the partitioning method that contains the most palindromic strings as the target partitioning method, and perform the partitioning of the string based on the target partitioning method to achieve deduplication of the string; wherein, the partitioning methods of different strings are different.

9. An electronic device, characterized in that, Including: A memory, configured to store a computer program; A processor, configured to implement the steps of the data processing method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the data processing method according to any one of claims 1 to 7 are implemented.