Data processing method, apparatus, device, medium, and product

By automatically determining the standardized form of the unit information entered by the user through the matching module, the problem of high management difficulty and high cost caused by non-standardized unit information is solved, and efficient information management and risk avoidance are achieved.

CN115344600BActive Publication Date: 2026-05-01CHINA CONSTRUCTION BANK +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA CONSTRUCTION BANK
Filing Date
2022-08-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, the organizational information filled in by users is not standardized, which leads to high difficulty and cost in information management and maintenance, as well as business and legal risks.

Method used

By responding to the unit information input by the user, the matching module matches it with unit information in multiple sets to determine the identifier of the target set, thereby automatically determining the standard unit information.

Benefits of technology

It reduces the difficulty and cost of information management and maintenance for organizations, improves the efficiency of information management, avoids related risks, and simplifies standardized information management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115344600B_ABST
    Figure CN115344600B_ABST
Patent Text Reader

Abstract

The application discloses a data processing method, device, equipment, medium and product. The method comprises the following steps: in response to a first input of a user to first unit information, determining the first unit information; matching the first unit information with unit information in a plurality of sets, each set comprising at least one unit information corresponding to one unit; in the case that the first unit information matches unit information in a target set, determining a target identifier corresponding to the target set as an identifier corresponding to the first unit information, each identifier being used for representing standard unit information of one unit, and the target set being one of the plurality of sets. By using the data processing method, device, equipment, medium and product provided by the application, the difficulty and cost of unit information management and maintenance can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and in particular relates to a data processing method, apparatus, device, medium and product. Background Technology

[0002] During the process of handling business, users are often required to fill in their organization information.

[0003] However, service providers cannot require users to learn standardized information filling techniques. Therefore, the organizational information filled in by users is often arbitrary and not standardized; it may contain abbreviations, alternative names, or even errors. This non-standardized information increases the difficulty and cost of information management and maintenance. Summary of the Invention

[0004] This application provides a data processing method, apparatus, device, medium, and product, which can at least solve the problems of high difficulty and high cost in information management and maintenance in the prior art.

[0005] In a first aspect, embodiments of this application provide a data processing method, the method comprising:

[0006] In response to the user's first input of the first unit information, determine the first unit information;

[0007] The first unit information is matched with the unit information in multiple sets, each set including at least one unit information corresponding to a unit;

[0008] If the information of the first unit matches the information of the unit in the target set, the target identifier corresponding to the target set is determined to be the identifier corresponding to the information of the first unit. Each identifier is used to represent the standard unit information of a unit, and the target set is one of multiple sets.

[0009] Secondly, embodiments of this application provide a data processing apparatus, the apparatus comprising:

[0010] The first determining module is used to determine the first unit information in response to the user's first input of the first unit information;

[0011] The matching module is used to match the first unit information with unit information in multiple sets, each set including at least one unit information corresponding to a unit;

[0012] The second determining module is used to determine the target identifier corresponding to the target set as the identifier corresponding to the first unit information when the first unit information matches the unit information in the target set. Each identifier is used to represent the standard unit information of a unit, and the target set is one of multiple sets.

[0013] Thirdly, embodiments of this application provide an electronic device, the device comprising: a processor and a memory storing computer program instructions;

[0014] When the processor executes the computer program instructions, it implements the data processing method as shown in any embodiment of the first aspect.

[0015] Fourthly, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the data processing method shown in any embodiment of the first aspect.

[0016] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform the data processing method shown in any embodiment of the first aspect.

[0017] The data processing method, apparatus, device, medium, and product of this application embodiment can, in response to a user's first input of first unit information, determine the first unit information, and then match the first unit information with unit information in multiple sets. If the first unit information matches unit information included in a target set among the multiple sets, the target identifier corresponding to the target set is determined to be the identifier corresponding to the first unit information. Since each set includes at least one unit information corresponding to a unit, if the first unit information successfully matches the unit information included in the target set, it can be determined that the unit corresponding to the first unit information and the unit corresponding to the unit information in the target set are the same unit. Therefore, the identifier corresponding to the first unit information can be determined as the identifier corresponding to the target set, i.e., the target identifier. Since each identifier is used to represent the standard unit information of a unit, determining the identifier corresponding to the first unit information determines the standard unit information corresponding to the first unit information. This allows for automatic determination of the corresponding standard unit information even when the user-input unit information is not standard, reducing the difficulty and cost of unit information management and maintenance. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart of a data processing method provided in one embodiment of this application;

[0020] Figure 2This is a flowchart of another data processing method provided in one embodiment of this application;

[0021] Figure 3 This is a schematic diagram of the structure of a data processing apparatus provided in one embodiment of this application;

[0022] Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation

[0023] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0024] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0025] Furthermore, it should be noted that the acquisition, storage, use, and processing of data in this application's technical solution all comply with the relevant provisions of national laws and regulations.

[0026] As in the background technology, during the process of handling business, users are often required to fill in company information, such as company name and company address.

[0027] However, service providers cannot require users to learn standardized information entry methods. Therefore, the company information entered by users is often arbitrary and non-standardized, potentially including abbreviations, alternative names, or even errors. This non-standardized information increases the difficulty and cost of information management and maintenance. Furthermore, non-standardized data can pose business and legal risks. Criminals can use fabricated or nearly legitimate company names and addresses to bypass risk control rules and obtain illegal resources.

[0028] Since it's impossible to require users to invest in learning standardized information filling techniques, the business processing party needs to standardize information management internally to mitigate potential risks. However, this type of data is used extensively and is linked to real-time transactions, placing stringent performance requirements; furthermore, from a business efficiency perspective, the system design should not be overly complex, thereby reducing efficiency.

[0029] Based on this, embodiments of this application provide a data processing method, apparatus, device, medium, and product that can, in response to a user's first input of first unit information, determine the first unit information, and then match the first unit information with unit information in multiple sets. If the first unit information matches unit information included in a target set among the multiple sets, the target identifier corresponding to the target set is determined to be the identifier corresponding to the first unit information. Since each set includes at least one unit information corresponding to a unit, when the first unit information successfully matches the unit information included in the target set, it can be determined that the unit corresponding to the first unit information and the unit corresponding to the unit information in the target set are the same unit. Therefore, the identifier corresponding to the first unit information can be determined as the identifier corresponding to the target set, i.e., the target identifier. Since each identifier is used to represent the standard unit information of a unit, determining the identifier corresponding to the first unit information determines the standard unit information corresponding to the first unit information. This allows for automatic determination of the corresponding standard unit information even when the user-input unit information is not standard, reducing the difficulty and cost of unit information management and maintenance.

[0030] Furthermore, by establishing a simple, flexible, and fast standardized management system for organizational information, it is possible to assist business recipients in managing and verifying organizational information and mitigating related risks, thereby playing a positive role in further improving the digitalization of business systems.

[0031] Figure 1 The diagram illustrates a flow chart of a data processing method according to an embodiment of this application. It should be noted that this data processing method can be applied to data processing devices, such as… Figure 1 As shown, the data processing method may include the following steps:

[0032] S110, in response to the user's first input of the first unit information, determine the first unit information;

[0033] S120, Match the first unit information with the unit information in multiple sets;

[0034] S130, if the first unit information matches the unit information in the target set, determine the target identifier corresponding to the target set as the identifier corresponding to the first unit information.

[0035] Therefore, in response to the user's initial input of first unit information, the first unit information can be determined. This first unit information is then matched against unit information in multiple sets. If the first unit information matches unit information included in a target set, the target identifier corresponding to that target set is determined as the identifier corresponding to the first unit information. Since each set includes at least one unit information corresponding to a unit, when the first unit information successfully matches unit information included in the target set, it can be determined that the unit corresponding to the first unit information and the unit corresponding to unit information in the target set are the same unit. Thus, the identifier corresponding to the first unit information can be determined as the identifier corresponding to the target set, i.e., the target identifier. Since each identifier represents the standard unit information of a unit, determining the identifier corresponding to the first unit information determines the standard unit information corresponding to the first unit information. This allows for automatic determination of the corresponding standard unit information even when the user's input unit information is non-standard, reducing the difficulty and cost of unit information management and maintenance.

[0036] Regarding S110, the first input can be the input for entering the information of the first unit.

[0037] In some implementations, the first unit information may include the first unit name and / or the first unit address, and the first unit address may include the first administrative division information and the first house number.

[0038] In some examples, users fill in the company name and address when conducting business.

[0039] In some implementations, to further reduce the likelihood of users entering incorrect unit addresses, if the first unit information includes a first unit address, the method may further include the following steps before determining the first unit information:

[0040] In response to the first input, determine the first address number included in the first unit address;

[0041] If the first address number is outside the preset range, the first address number will be cleared and a prompt message will be displayed to remind the user to re-enter an address number that is within the preset range.

[0042] Here, the preset range can be a pre-set range of actual house numbers. If the first house number entered by the user exceeds the preset range, it indicates that the first house number is an incorrect house number. The data processing device can automatically clear the first house number and display a prompt message, prompting the user that the house number was entered incorrectly and that a house number belonging to the preset range needs to be re-entered.

[0043] In some examples, the preset range for house numbers is greater than 0 and less than or equal to 100. When the user enters the first house number, the data processing device automatically clears the first house number "101" and displays the prompt message "The house number is incorrect. Please enter a house number that is greater than 0 and less than or equal to 100".

[0044] Therefore, by setting a preset range for house numbers to restrict the input of house numbers, the possibility of users entering incorrect unit addresses can be greatly reduced.

[0045] Regarding S120, there is a one-to-one correspondence between sets and units. Each set can include at least one unit information corresponding to a unit. Specifically, each set can include at least one unit name or unit address corresponding to a unit. That is, the at least one unit name is a different name for a unit, such as a standard unit name and non-standard unit names such as abbreviations or aliases; the at least one unit address is a different representation of a unit's address, such as a standard unit address and non-standard unit addresses such as abbreviations or aliases. By matching the first unit information with the unit information in multiple sets, the set to which the unit information matching the first unit information belongs can be determined.

[0046] Regarding S130, the target set can be one of multiple sets. If the unit information in the target set matches the first unit information, it indicates that the first unit information and the unit information in the target set belong to the same unit. Therefore, the target identifier corresponding to the target set can be determined as the identifier corresponding to the first unit information. Here, each set can correspond to a unique identifier, and each identifier can be used to represent the standard unit information of a unit. Therefore, the first unit information can be managed and maintained based on this target identifier. In this way, even if the first unit information entered by the user is not standardized, it will not cause inconvenience to the management and maintenance of the information.

[0047] In some implementations, to enrich the unit information in the set and thus more accurately identify the unit information input by the user, the method may further include, if the first unit information matches the unit information in the target set:

[0048] If the information of the first unit is not repeated with any information of any unit in the target set, add the information of the first unit to the target set.

[0049] Here, if the first unit information matches the unit information in the target set, but does not overlap with any unit information in the target set, then the first unit information can be added to the target set to enrich the unit information in the target set.

[0050] Specifically, a similarity threshold can be set. If the similarity between the first unit information and the unit information in the target set is greater than the similarity threshold, it can be determined that the first unit information matches the unit information in the target set. If the similarity between the first unit information and the unit information in the target set is greater than the similarity threshold but less than 100%, the first unit information can be added to the target set.

[0051] In some examples, the similarity threshold can be 80%. If the similarity between the first unit information and the unit information in the target set is greater than 80% but less than 100%, the first unit information can be added to the target set.

[0052] In this way, adding unit information that matches but does not duplicate the unit information in the target set to the target set can enrich the unit information in the set, thereby making it easier to more accurately identify the unit information input by the user.

[0053] In some implementations, to more accurately identify the unit information input by the user, the method may further include, after S120:

[0054] If the information of the first unit does not match the information of any of the multiple sets, a first set is established based on the information of the first unit.

[0055] Here, if the information of the first unit does not match the information of any of the multiple sets, it means that there is no set corresponding to the unit that corresponds to the information of the first unit in the established set. Therefore, it is necessary to establish a set corresponding to the unit.

[0056] In some examples, there are already established sets A, B, and C. The first unit information 'a' input by the user does not match the unit information in sets A, B, and C. Therefore, set D can be established, which includes the first unit information 'a'. At the same time, the identifier corresponding to set D can also be generated.

[0057] This increases the number and richness of sets, making it easier to find sets that correspond to the unit information input by the user, thus improving the accuracy of recognizing the unit information input by the user.

[0058] In some embodiments, to improve the accuracy of unit information matching, such as Figure 2As shown, S120 may include S1201-S1203, and S130 may include S1301, wherein:

[0059] S1201, perform word segmentation on the first unit information to obtain multiple words corresponding to the first unit information.

[0060] Here, we can first use a word segmenter to segment the first unit information to obtain multiple words corresponding to the first unit information.

[0061] S1202, filter the preset word segments from multiple word segments to obtain at least one first word segment.

[0062] Here, a filter can be used to filter out preset word segments from multiple word segments. These preset word segments can include pre-set invalid words, punctuation marks, and / or other special symbols.

[0063] In some implementations, to manage the matching process of unit information more flexibly and improve the accuracy of unit information matching, the method may further include the following before S1202:

[0064] In response to a second input from the user identifying the target filter, the target filter is determined;

[0065] In response to a third user input regarding the preset word segmentation, a target filter is set to filter the preset word segmentation.

[0066] Here, multiple filters can be preset, such as word segmentation filters and character filters. The word segmentation filter can be used to filter words, while the character filter can be used to filter letters, numbers, punctuation marks, and / or other special symbols. Target filters can be invoked by importing class files.

[0067] Specifically, the second input can be the input from which the user selects a target filter from a set of preset filters, and the third input can be the input from which preset word segments are set. The user can determine the target filter through the second input, and then set the word segments to be filtered by the target filter, i.e., the preset word segments, through the third input. Then, the user can use the target filter to filter the preset word segments from multiple word segments to obtain at least one first word segment.

[0068] In some examples, if a user selects a word segmentation filter as the target filter and then sets "company" as the preset word segmentation corresponding to the target filter, the target filter can filter out "company" from multiple word segments; if a user selects a character filter as the target filter and then sets "," as the preset word segmentation corresponding to the target filter, the target filter can filter out "," from multiple word segments.

[0069] In this way, users can flexibly select filters and set preset words to be filtered out according to their needs, thereby improving the accuracy of unit information matching.

[0070] S1203, Match the first word segment with the feature word segment corresponding to the feature unit information in each set.

[0071] Here, the feature unit information can be selected from the unit information included in the set. The set usually contains a large number of unit information entries, and matching each one with the first unit information would be very time-consuming. Therefore, a small number of unit information entries from each set can be selected as the feature unit information for that set, and the first unit information is matched against the feature unit information of each set. To improve matching efficiency, the number of feature unit information entries can be no more than four. Specifically, during matching, the first word corresponding to the first unit information is matched against the feature word corresponding to the feature unit information in each set. Each unit information entry in each set is stored in the form of word segments. For example, it can store the set type (unit name set, unit address set, other sets), unit information, the frequency of use of the unit information, each word corresponding to the unit information, and the hash value of each word.

[0072] In some implementations, to make the selected feature unit information in each set more representative, the method may further include, before S1203:

[0073] The characteristic unit information of each set is determined by at least one of the following methods:

[0074] Among the various unit information included in the set, the feature unit information that is entered most frequently by the user is the unit information in the set;

[0075] Among the multiple unit information included in the set, the unit information with the smallest dispersion of difference from other unit information in the set is the characteristic unit information in the set;

[0076] In response to the user's fourth input of preset unit information, the preset unit information is determined to be the feature unit information in the set.

[0077] Here, the unit information with the most user inputs in the set can be used as the feature unit information of the set; the difference between each unit information in the set and other unit information in the set can be calculated, and the unit information with the smallest dispersion of the difference between the unit information and other unit information can be used as the feature unit information of the set; or a preset unit information can be directly set by the staff as the feature unit information of the set.

[0078] In addition, it can be determined that among the multiple units of information included in the set, those with more than one number of user inputs are characteristic units of information in the set; or it can be determined that among the multiple units of information included in the set, those with a dispersion of differences from other units of information in the set that exceeds a dispersion threshold are characteristic units of information in the set.

[0079] In this way, the selected feature unit information in each set can be more representative, improving the accuracy of unit information matching.

[0080] S1301, if the first word segmentation matches the target feature word segmentation corresponding to the target set, determine the target identifier corresponding to the target set as the identifier corresponding to the first unit information.

[0081] Thus, by matching the first word corresponding to the first unit information with the word corresponding to the feature unit information in each set through the above process, the accuracy of unit information matching can be further improved.

[0082] In related technologies, unit information data is maintained in Redis. Once the Redis data is lost, all data is lost. Furthermore, Redis data initialization takes a lot of time, which is not conducive to the rapid deployment of the system.

[0083] Based on this, in some implementations, to avoid data loss and improve query efficiency, the method may further include:

[0084] Instantiate multiple collections of data into an Oracle database and cache them in Redis.

[0085] Specifically, by using a data layer primarily based on Oracle and secondarily based on Redis to store various types of data required for the data processing device to run, including multiple sets of data, and storing frequently read and written critical data in Redis, the burden on the Oracle database can be reduced, query efficiency can be improved, and since the data has already been instantiated in the Oracle database, it can be recovered even if the Redis data is lost.

[0086] In related technologies, data maintenance can only be achieved by modifying the database or Redis cache, which is very inconvenient. For example, administrative division data is maintained in a database, requiring version control of SQL for modification and maintenance.

[0087] Therefore, in some implementations, to facilitate data maintenance, the method may also provide a data maintenance page to maintain the data stored in the database.

[0088] Here, for maintaining unit addresses, you can provide pages for maintaining administrative divisions, address numbers, and road information, etc.

[0089] Specifically, administrative division data maintenance can include: provincial-level administrative division maintenance (provinces, autonomous regions, municipalities, special administrative regions), prefecture-level administrative division maintenance (prefecture-level cities, regions, autonomous prefectures, leagues), county-level administrative division maintenance (municipal districts, county-level cities, counties, self-governed counties, banners, autonomous banners, special zones, forest areas), and township-level administrative division maintenance (streets, towns, townships, ethnic townships, sumu, ethnic sumu, county-level districts). The administrative division maintenance page can include basic functions such as adding, deleting, modifying, and querying data, as well as importing and exporting (for rare administrative division changes).

[0090] The system maintains address and road information, supports functions such as restricting address ranges, and further strengthens the screening of false data.

[0091] In this way, by providing data maintenance pages such as administrative division maintenance pages and address and road information maintenance pages, users can maintain the data in the database more flexibly.

[0092] Generally, after a new set is created based on a certain unit information, if new unit information that matches that unit information is obtained, a new set will not be created again. However, in some special scenarios, set classification may occur. For example, set A is created based on unit information 'a', and then unit information 'b' and unit information 'c' are obtained. It is found that unit information 'b' and unit information 'c' match. However, when the program makes a judgment, it finds that unit information 'b' matches unit information 'a', so a new set is not created. But unit information 'a' and unit information 'c' do not match, so a new set B is created based on unit information 'c', resulting in set splitting.

[0093] In some implementations, a collection maintenance function can be built to prevent data accuracy issues caused by collection splitting. Users can maintain collections and unit information within collections through the collection maintenance page. The entry point for the collection maintenance function can be found on the address information maintenance page (including the administrative division maintenance page and the house number and road information maintenance page). The collection maintenance page provides the function of searching for collections by characteristic unit information or unit information, as well as the function of modifying collections. Users can perform operations such as merging and splitting collections according to the actual situation.

[0094] The relevant technologies cannot intuitively display the hierarchy and affiliation of administrative divisions, which is not conducive to flexibly responding to business scenarios.

[0095] Therefore, in some implementations, to more intuitively display the relationships between different levels of administrative divisions, the administrative division maintenance page in the data processing method provided in this application embodiment can use both map display and tree diagram display modes. This allows for a more intuitive display of the relationships between different levels of administrative divisions. Data initialization can rely on existing data.

[0096] In related technologies, when a batch of unit information is concentrated at the same address, if this unit information does not match the unit information in the existing set, a new set needs to be created. At this time, to prevent the generation of conflicting sets, the program locks the set in that region until a new set is created. During this process, this batch of unit information will be in a waiting state, causing thread congestion and other problems. Furthermore, the program uses loops to compare and match data item by item, resulting in poor performance.

[0097] Based on this, in some implementations, a query tree can be constructed for each administrative region based on the hash value array of unit information, thereby improving matching efficiency. In this case, if a piece of data needs to be identified, it is segmented by a tokenizer and then queried. If matching unit information is found, the identifier of the set to which that unit information belongs is returned; otherwise, a new set is created. To prevent large-scale thread waiting caused by locking the set, the locking scope can be adjusted to the tree nodes of the query tree, minimizing the impact and improving program efficiency. For example, the locking scope can be changed from the prefecture-level administrative region to the county-level administrative region.

[0098] Based on the same inventive concept, embodiments of this application also provide a data processing system, which can be divided into a visual layer, a service layer, and a data layer.

[0099] The visual layer provides various maintenance pages, such as the address information maintenance page and the word segmenter maintenance page, for data maintenance of the overall data processing system.

[0100] The service layer provides encapsulated services for collections, queries, and word segmentation, offering services to the front-end and other developers. Examples include the front-end's query function and word segmentation queries and collection queries provided to other developers.

[0101] The data layer is primarily based on Oracle and secondarily on Redis. It stores various data required for system operation, while Redis stores frequently read and written critical data, reducing the burden on the Oracle database and improving query efficiency.

[0102] Based on the same inventive concept, embodiments of this application also provide a data processing apparatus. The following, in conjunction with… Figure 3 The data processing apparatus provided in the embodiments of this application will be described in detail.

[0103] Figure 3 A schematic diagram of the structure of a data processing apparatus provided in one embodiment of this application is shown.

[0104] like Figure 3 As shown, the data processing apparatus may include:

[0105] The first determining module 301 is used to determine the first unit information in response to the user's first input on the first unit information;

[0106] The matching module 302 is used to match the first unit information with unit information in multiple sets, each set including at least one unit information corresponding to a unit;

[0107] The second determining module 303 is used to determine the target identifier corresponding to the target set as the identifier corresponding to the first unit information when the first unit information matches the unit information in the target set. Each identifier is used to represent the standard unit information of a unit, and the target set is one of multiple sets.

[0108] Therefore, in response to the user's initial input of first unit information, the first unit information can be determined. This first unit information is then matched against unit information in multiple sets. If the first unit information matches unit information included in a target set, the target identifier corresponding to that target set is determined as the identifier corresponding to the first unit information. Since each set includes at least one unit information corresponding to a unit, when the first unit information successfully matches unit information included in the target set, it can be determined that the unit corresponding to the first unit information and the unit corresponding to unit information in the target set are the same unit. Thus, the identifier corresponding to the first unit information can be determined as the identifier corresponding to the target set, i.e., the target identifier. Since each identifier represents the standard unit information of a unit, determining the identifier corresponding to the first unit information determines the standard unit information corresponding to the first unit information. This allows for automatic determination of the corresponding standard unit information even when the user's input unit information is non-standard, reducing the difficulty and cost of unit information management and maintenance.

[0109] In some implementations, to enrich the unit information in the set and thus more accurately identify the unit information input by the user, the device may further include, if the first unit information matches the unit information in the target set:

[0110] The add module is used to add the first unit information to the target set when the first unit information does not overlap with any unit information in the target set.

[0111] In some embodiments, to more accurately identify the unit information input by the user, the device may further include:

[0112] A module is established to create a first set based on the first unit information after matching the first unit information with the unit information in multiple sets, in the case where the first unit information does not match the unit information included in any of the multiple sets.

[0113] In some implementations, to improve the accuracy of unit information matching, the matching module 302 may specifically include:

[0114] The word segmentation submodule is used to segment the first unit information to obtain multiple words corresponding to the first unit information.

[0115] The filtering submodule is used to filter preset word segments from multiple word segments to obtain at least one first word segment.

[0116] The matching submodule is used to match the first word segment with the feature word segment corresponding to the feature unit information in each set. The feature unit information is selected from the unit information included in the set.

[0117] Based on this, the second determining module 303 may specifically include:

[0118] The determination submodule is used to determine the target identifier corresponding to the target set as the identifier corresponding to the first unit information when the first word segmentation matches the target feature word segmentation corresponding to the target set.

[0119] In some embodiments, to more flexibly manage the matching process of unit information and improve the accuracy of unit information matching, the device may further include:

[0120] The third determining module is used to determine the target filter in response to a second input from the user regarding the identifier of the target filter before filtering preset words from multiple word segments to obtain at least one first word.

[0121] The settings module is used to respond to the user's third input on the preset word segmentation, and the target filter is used to filter the preset word segmentation.

[0122] In some embodiments, to make the selected feature unit information in each set more representative, the device may further include:

[0123] The fourth determining module is used to determine the feature unit information in each set by at least one of the following methods before matching the first word segment with the feature word segment corresponding to the feature unit information in each set:

[0124] Among the various unit information included in the set, the feature unit information that is entered most frequently by the user is the unit information in the set;

[0125] Among the multiple unit information included in the set, the unit information with the smallest dispersion of difference from other unit information in the set is the characteristic unit information in the set;

[0126] In response to the user's fourth input of preset unit information, the preset unit information is determined to be the feature unit information in the set.

[0127] In some implementations, the first unit information includes the first unit name and / or the first unit address, whereby the first unit address includes first administrative division information and first house number.

[0128] In some embodiments, to further reduce the likelihood of users entering incorrect unit addresses, the device may further include:

[0129] The fifth determining module is used to determine, in response to the first input, the first address number included in the first unit address before determining the first unit information, if the first unit information includes the first unit address;

[0130] The processing module is used to clear the first address number and display a prompt message when the first address number exceeds the preset range, so as to prompt the user to re-enter an address number that belongs to the preset range.

[0131] Figure 4 A schematic diagram of the structure of an electronic device provided in one embodiment of this application is shown.

[0132] like Figure 4 As shown, the electronic device 4 is a structural diagram of an exemplary hardware architecture of an electronic device that can implement the data processing method and data processing apparatus according to the embodiments of this application. This electronic device may refer to the electronic device in the embodiments of this application.

[0133] The electronic device 4 may include a processor 401 and a memory 402 storing computer program instructions.

[0134] Specifically, the processor 401 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0135] Memory 402 may include a large-capacity memory for data or instructions. For example, and not limitingly, memory 402 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be internal or external to an integrated gateway disaster recovery device. In a particular embodiment, memory 402 is non-volatile solid-state memory. In a particular embodiment, memory 402 may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. Thus, generally, memory 402 includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of this application.

[0136] The processor 401 implements any of the data processing methods described in the above embodiments by reading and executing computer program instructions stored in the memory 402.

[0137] In one example, the electronic device may also include a communication interface 403 and a bus 404. Wherein, as... Figure 4 As shown, the processor 401, memory 402, and communication interface 403 are connected through bus 404 and complete communication with each other.

[0138] The communication interface 403 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0139] Bus 404 includes hardware, software, or both, that couples components of an electronic device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 404 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.

[0140] The electronic device can execute the data processing method described in the embodiments of this application, thereby achieving the combination Figures 1 to 3 The data processing methods and apparatus described.

[0141] Furthermore, in conjunction with the data processing methods in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the data processing methods in the above embodiments.

[0142] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0143] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0144] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0145] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0146] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A data processing method, characterized in that, The method includes: In response to the user's first input of the first unit information, the first unit information is determined; The first unit information is matched with unit information in multiple sets, each set including at least one unit information corresponding to a unit; If the first unit information matches the unit information in the target set, the target identifier corresponding to the target set is determined to be the identifier corresponding to the first unit information. Each identifier is used to represent the standard unit information of a unit. The target set is one of the plurality of sets. Matching the first unit information with unit information in multiple sets, including: The first unit information is segmented into words to obtain multiple words corresponding to the first unit information. Filter the preset word segments from the multiple word segments to obtain at least one first word segment; The first word segment is matched with the feature word segment corresponding to the feature unit information in each of the sets, wherein the feature unit information is selected from the unit information included in the set; Before filtering the preset word segments from the plurality of word segments to obtain at least one first word segment, the method further includes: In response to a second input from the user identifying the target filter, the target filter is determined, the target filter including a word segmentation filter and / or a character filter; In response to a third user input on preset word segments, a target filter is set to filter the preset word segments, which include pre-set invalid words, punctuation marks, and / or special characters.

2. The method according to claim 1, characterized in that, If the first unit information matches the unit information in the target set, the method further includes: If the first unit information is not duplicated with any unit information in the target set, the first unit information is added to the target set.

3. The method according to claim 1, characterized in that, After matching the first unit information with unit information in multiple sets, the method further includes: If the first unit information does not match the unit information included in any of the plurality of sets, a first set is established based on the first unit information.

4. The method according to claim 1, characterized in that, The step of determining the target identifier corresponding to the target set as the identifier corresponding to the first unit information when the first unit information matches the unit information in the target set includes: If the first word segmentation matches the target feature word segmentation corresponding to the target set, the target identifier corresponding to the target set is determined to be the identifier corresponding to the first unit information.

5. The method according to claim 4, characterized in that, Before matching the first word segment with the feature word segment corresponding to the feature unit information in each of the sets, the method includes: The characteristic unit information in each set is determined by at least one of the following methods: Among the multiple unit information included in the set, the feature unit information that is entered most frequently by the user is the one in the set; Among the multiple unit information included in the set, the unit information with the smallest dispersion of difference from other unit information in the set is identified as the feature unit information in the set. In response to the user's fourth input of preset unit information, the preset unit information is determined to be the feature unit information in the set.

6. The method according to claim 1, characterized in that, The first unit information includes the first unit name and / or the first unit address, and the first unit address includes the first administrative division information and the first house number.

7. The method according to claim 6, characterized in that, If the first unit information includes a first unit address, the method further includes, before determining the first unit information: In response to the first input, determine the first address number included in the first unit address; If the first address number is outside the preset range, the first address number is cleared and a prompt message is displayed to prompt the user to re-enter an address number that is within the preset range.

8. A data processing apparatus, characterized in that, The device includes: The first determining module is used to determine the first unit information in response to the user's first input on the first unit information; The matching module is used to match the first unit information with unit information in multiple sets, each set including at least one unit information corresponding to a unit; The second determining module is used to determine the target identifier corresponding to the target set as the identifier corresponding to the first unit information when the first unit information matches the unit information in the target set. Each identifier is used to represent the standard unit information of a unit, and the target set is one of the plurality of sets. The matching module includes: The word segmentation submodule is used to perform word segmentation processing on the first unit information to obtain multiple words corresponding to the first unit information; The filtering submodule is used to filter preset words from the multiple word segments to obtain at least one first word; The matching submodule is used to match the first word segment with the feature word segment corresponding to the feature unit information in each of the sets, wherein the feature unit information is selected from the unit information included in the set; The device further includes: The third determining module is used to determine the target filter in response to a second input from the user regarding the identifier of the target filter before filtering the preset word segments in the plurality of word segments to obtain at least one first word segment. The target filter includes a word segmentation filter and / or a character filter. The setting module is used to respond to a third user input on preset word segments and set the target filter to filter the preset word segments. The preset word segments include pre-set invalid words, punctuation marks and / or special symbols.

9. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the data processing method as described in any one of claims 1-7.

10. A computer storage medium, characterized in that, The computer storage medium stores computer program instructions, which, when executed by a processor, implement the data processing method as described in any one of claims 1-7.

11. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the data processing method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Address recognition method and device

    CN110765280A