Method, apparatus, and program product for managing data schema

By dividing the acquisition devices into groups and determining the shared data pattern, the problem of transmission and storage waste caused by repetitive patterns in the data is solved, and more efficient data management is achieved.

CN114861743BActive Publication Date: 2025-12-16EMC IP HLDG CO LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110075083.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-20
Publication Date
2025-12-16
Estimated Expiration
2041-01-20

AI Technical Summary

Technical Problem

In existing technologies, data from acquisition devices contains duplicate data patterns, which leads to a waste of transmission bandwidth and storage space, and the performance of removing duplicate data at edge computing devices is not ideal.

Method used

Multiple acquisition devices are divided into multiple groups, and a shared data mode is determined for each group. Duplicate data is removed before transmission by edge computing devices, reducing the transmission bandwidth requirement.

Benefits of technology

Effectively manage data patterns to reduce resource overhead in data storage and transmission, and improve data transmission and storage efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114861743B_ABST
    Figure CN114861743B_ABST
Patent Text Reader

Abstract

The present disclosure relates to methods, devices and program products for managing data patterns. In one method, a plurality of sets of data patterns respectively associated with a plurality of collection devices are obtained, one of the plurality of sets of data patterns representing a pattern of duplicated data in data from one of the plurality of collection devices. Based on clustering of the plurality of sets of data patterns, the plurality of collection devices are divided into a plurality of groups. Based on the sets of data patterns respectively associated with the collection devices in one of the plurality of groups, a set of shared data patterns for sharing among the collection devices in the group is determined. Corresponding devices and computer program products are provided. With exemplary implementations of the present disclosure, data patterns that can be shared among the plurality of collection devices can be determined in a more accurate and efficient manner, thereby facilitating removal of duplicated data from the plurality of collection devices.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Implementations of the present disclosure relate to data processing, and more particularly, to a method, device and computer program product for managing data patterns of data from collection devices. BACKGROUND

[0002] With the development of data collection technology and network technology, various types of collection devices can be deployed in various application environments to monitor various parameters of the application environments. For example, in a road management environment, collection devices can be deployed at various geographical locations in a road to monitor the road status. In a production control environment, collection devices can be deployed at various locations of a production line to monitor the operation of the production line, etc. The data from the collection devices can be transmitted to a data center for later processing and analysis. It will be appreciated that there can be repetitive data (hereinafter, the pattern of repetitive data is referred to as a data pattern) in the collected data, which will cause waste of transmission bandwidth and storage space. At this time, how to more effectively manage the data patterns of the collected data becomes a research hotspot. SUMMARY

[0003] Therefore, it is desirable to develop and implement a technical solution to manage data patterns in a more effective manner. It is desirable that the technical solution can manage data patterns of data from multiple different collection devices in a more convenient and effective manner. Further, it is desirable that various data patterns can be utilized to improve the transmission and storage efficiency of data.

[0004] According to a first aspect of the present disclosure, a method for managing data patterns is provided. In the method, a plurality of sets of data patterns respectively associated with a plurality of collection devices are obtained, and a set of data patterns in the plurality of sets of data patterns represents a pattern of repetitive data in data from one of the plurality of collection devices. Based on clustering of the plurality of sets of data patterns, the plurality of collection devices are divided into a plurality of groups. Based on the sets of data patterns associated with the collection devices in a group of the plurality of groups, a set of shared data patterns for sharing among the collection devices in the group is determined.

[0005] According to a second aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; a volatile memory; and a memory coupled with the at least one processor, the memory having stored therein instructions that, when executed by the at least one processor, cause the device to perform the method according to the first aspect of the present disclosure.

[0006] According to a third aspect of the disclosure, there is provided a computer program product tangibly stored on a non-transitory computer readable medium and comprising machine executable instructions for performing the method according to the first aspect of the disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0007] The features, advantages, and other aspects of the present disclosure will become more apparent from the following detailed description in conjunction with the accompanying drawings, in which several implementations of the present disclosure are illustrated. The detailed description and drawings are merely illustrative of the present disclosure. The present disclosure is not limited to the implementations described and depicted herein.

[0008] Figure 1 A block diagram schematically illustrating an application environment in which an example implementation of the present disclosure can be implemented is shown;

[0009] Figure 2 A block diagram schematically illustrating a process for managing data schemas according to an example implementation of the present disclosure is shown;

[0010] Figure 3 A flow diagram schematically illustrating a method for managing data schemas according to an example implementation of the present disclosure is shown;

[0011] Figure 4 A block diagram schematically illustrating a process for obtaining a set of data schemas associated with one acquisition device according to an example implementation of the present disclosure is shown;

[0012] Figure 5 A block diagram schematically illustrating a data structure of schema features according to an example implementation of the present disclosure is shown;

[0013] Figure 6 A block diagram schematically illustrating a process for obtaining a set of shared data schemas according to an example implementation of the present disclosure is shown;

[0014] Figure 7 A block diagram schematically illustrating a deployment of a set of shared data schemas according to an example implementation of the present disclosure is shown;

[0015] Figure 8 A block diagram schematically illustrating a process for generating deduplicated data according to an example implementation of the present disclosure is shown; and

[0016] Figure 9 A block diagram schematically illustrating a device for managing data schemas according to an example implementation of the present disclosure is shown. DETAILED DESCRIPTION

[0017] Preferred implementations of the present disclosure will be described herein below with reference to a few preferred implementations. While the preferred implementations of the present disclosure are illustrated with reference to the attached figures, it is to be understood that the present disclosure can be embodied in various forms and should not be limited to the implementations set forth herein. Rather, these implementations are provided as example implementations of the present disclosure to convey the scope of the present disclosure to those skilled in the art.

[0018] The term "includes" and its variants are used inclusively, i.e., "includes but is not limited to." The term "or" means "and / or" unless otherwise specified. The term "based on" means "based, at least in part, on." The terms "one example implementation" and "an implementation" mean "at least one example implementation." The term "another implementation" means "at least one additional implementation." The terms "first," "second," and the like can refer to different or the same objects. Other explicitly and implicitly recited definitions can also be included below.

[0019] For ease of description, first refer to Figure 1 An application environment according to one example implementation of the present disclosure is described. The application environment can include an edge network and a core network. Figure 1 A block diagram of an application environment 100 in which example implementations of the present disclosure can be implemented is schematically shown. For example, an application environment 100 of the Internet of Things can be implemented in which an example implementation of the present disclosure can be implemented. The application environment 100 can include two parts, an edge network 110 and a core network 120. The edge network 110 can include an edge computing device 114 and a plurality of collection devices 112-1, 112-2, 112-3, 112-4, 112-5, and 112-6, etc. (collectively referred to as collection devices 112) connected to the edge computing device 114. The collection devices 112 can be of various types, such as cameras, video cameras, facsimile machines, scanners, etc. The edge computing device 114 can transmit data from the various collection devices 112 to server devices 122,..., and 124, etc. in the core network 120 for storage and subsequent further processing.

[0020] It will be appreciated that there can be repetitive content in the data from the collection devices 112. Techniques have been proposed to extract data patterns from the collected data. A data pattern library can be predefined, and when a data segment in the collected data is found to match a data pattern in the data pattern library, an identifier of the data pattern can be used to replace the data segment. A variety of data patterns 130 can be deployed at the core network 120. The repetitive content in the received data can be removed based on the variety of data patterns 130. In this way, the data can be stored in a more efficient manner.

[0021] Although the above technical solution can improve the storage efficiency at the core network 120, the data transmitted from the edge network 110 to the core network 120 is the original data including the repeated contents, and the transmission process will occupy a high bandwidth. In another technical solution, a data pattern for removing the repetition can be deployed in the edge computing device 114. However, there can be a large number (e.g., tens of thousands or even more) of data patterns, and due to the limitation of the storage space and the computing resources of the edge computing device 114, the performance of removing the repeated data at the edge computing device 114 is not ideal.

[0022] To solve the above-mentioned defects, the implementation manners of the present disclosure provide a technical solution for managing data patterns. According to one exemplary implementation manner of the present disclosure, a technical solution of classifying a plurality of collection devices 112 is proposed. In the following, referring to Figure 2 A summary according to one exemplary implementation manner of the present disclosure is described. Figure 2 A block diagram of a process 200 for managing data patterns according to an exemplary implementation manner of the present disclosure is schematically shown. As Figure 2 shown, a plurality of collection devices 112 in the edge network 110 can be divided into different groups. For example, a group of collection devices 210 can include the collection devices 112-1, 112-2, etc., a group of collection devices 220 can include the collection devices 112-4, 112-6, etc., and a group of collection devices 230 can include the collection devices 112-3, 112-5, etc.

[0023] A respective shared data pattern can be determined for each group of collection devices, for example, the group of collection devices 210 can have a shared data pattern 212. The shared data pattern herein refers to that the data from part and / or all of the collection devices 112-1, 112-2, etc. in the group of collection devices 210 includes the shared data pattern 212. According to one exemplary implementation manner of the present disclosure, the shared data pattern 212 can include one or more data patterns. As Figure 2 shown, the group of collection devices 220 can have a shared data pattern 222, and the group of collection devices 230 can have a shared data pattern 232. With the exemplary implementation manner of the present disclosure, a large number of data patterns do not need to be maintained at the edge computing device 114 at the edge network 110, but only the shared data pattern for a group of collection devices can be maintained. In this way, the burden of the edge computing device 114 can be reduced, and thus the edge computing device 114 can remove the repeated part from the collection device 112 before transmitting the data, thereby reducing the demand for transmission bandwidth.

[0024] In the following, referring to Figure 3 more details according to one exemplary implementation manner of the present disclosure are described in detail. Figure 3A flowchart of a method 300 for managing data patterns according to an example implementation of the present disclosure is schematically shown. At block 310, a plurality of sets of data patterns respectively associated with a plurality of collection devices are obtained. Here, a set of data patterns in the plurality of sets of data patterns represents a pattern of repeated data in data from one of the plurality of collection devices.

[0025] In the following, reference will be made to Figure 4 The process of obtaining a set of data patterns from one collection device 112 can be performed in a similar manner to obtain a plurality of sets of data patterns respectively from a plurality of collection devices. Figure 4 A block diagram of a process 400 for obtaining a set of data patterns associated with one collection device according to an example implementation of the present disclosure is schematically shown. At Figure 4 In the following, reference will be made to

[0026] According to one example implementation of the present disclosure, a set of initial data patterns 420 associated with the collection device 112 can be extracted as indicated by arrow 442. Specifically, data patterns occurring in the data 420 can be obtained based on a variety of pattern extraction techniques that have been developed and / or will be developed in the future. For example, data patterns 422, 424,..., 426,..., and 428 can be obtained. Specifically, a predetermined time length can be specified. For example, data 410 in the past one month (or other time length) can be obtained, and the number of occurrences of each data pattern in the past one month can be counted.

[0027] The set of initial data patterns 420 can be sorted based on the number of occurrences of each data pattern in the set of initial data patterns 420 in the data 410 as indicated by arrow 444. Assuming that q data patterns are collected, the top n data patterns can be selected from the sorted plurality of data patterns. As Figure 4 indicated, data patterns 424,..., and 426, etc. can be selected. According to one example implementation of the present disclosure, similar operations can be performed for data from each collection device, and the top set of n data patterns can be obtained respectively from each collection device 112.

[0028] In the following, reference will be made to Figure 3How to process the collected multiple sets of data patterns is described. At block 320, the multiple collection devices are divided into multiple groups based on clustering of the multiple sets of data patterns. It will be appreciated that each set of data patterns can include n data patterns, and assuming there are m collection devices, n data patterns can be obtained for each collection device from the m collection devices. According to one example implementation of the present disclosure, a unique identifier can be utilized to represent each data pattern. For example, the number of bits p of the identifier can be pre-specified, and an integer between 0 and 2 p

[0029] According to one example implementation of the present disclosure, multiple pattern features of the multiple sets of data patterns can be generated based on the multiple sets of data patterns, respectively. In the following, more details about the pattern features are described. Figure 5 Figure 5 A block diagram of a data structure 500 of a pattern feature according to an example implementation of the present disclosure is schematically shown. The pattern feature as shown in Figure 5 may be generated for each set of data patterns. As shown in Figure 5 For a set of data patterns from a given collection device 112, the pattern feature can be generated based on the identifiers (IDs) of the individual data patterns included in the set of data patterns. Specifically, the pattern feature 510 can include data pattern ID 512, …, and data pattern ID 514. At this time, the pattern feature can include the identifiers of n data patterns.

[0030] It will be appreciated that the number of occurrences of the individual data patterns in the data is different, and for the data patterns with higher number of occurrences, it is more helpful to reduce the amount of data based on the data patterns to remove the duplicated data. Thus, the pattern feature can be determined based on the number of occurrences of the data patterns. According to one example implementation of the present disclosure, for a set of data patterns, the pattern feature of the set of data patterns can be generated based on the individual data patterns in the set of data patterns and the number of occurrences of the individual data patterns in the data. As shown in Figure 5 the data pattern 510 can further include the number of occurrences 522, …, and 524 of the individual data patterns.

[0031] It will be appreciated that Figure 5 Only one data structure of the pattern feature 510 is schematically shown. According to one example implementation of the present disclosure, the pattern feature can be generated based on other manners. For example, the pattern feature 510 can include multiple fields, each field can correspond to one data pattern, and the number in each field can represent the number of occurrences of the data pattern. According to one example implementation of the present disclosure, the pattern feature of a set of data patterns can also be determined based on other manners. ​​

[0032] Further, the plurality of collection devices can be divided into a plurality of groups based on the plurality of pattern features. It will be appreciated that the pattern features herein can have a high dimensionality and be difficult to cluster, and thus a dimensionality reduction process can be performed on the pattern features and the reduced dimensionality features can be clustered. According to one example implementation of the present disclosure, the plurality of pattern features can be converted to a plurality of low dimensionality features, for example, based on any of a Multidimensional Scaling (MDS) algorithm, a Principal Component Analysis (PCA) algorithm, a Linear Discriminant Analysis (LDA) algorithm, a Local Linear Embedding (LLE) algorithm, etc. Here, each of the low dimensionality features has a lower dimensionality than the original pattern features.

[0033] According to one example implementation of the present disclosure, the high dimensionality pattern features can be reduced to two dimensions or other lower dimensionality. In turn, the plurality of low dimensionality features can be divided into a plurality of clusters based on distances between the low dimensionality features. Here, each cluster can include one or more low dimensionality features, and each low dimensionality feature is associated with a collection device. In this manner, the original plurality of collection devices can be divided into a plurality of groups.

[0034] Having described dividing the plurality of collection devices into a plurality of groups, in the following, a process for determining a set of shared data patterns for each of the plurality of groups will be described with reference to Figure 3 Continuing with the description of determining a set of shared data patterns for each of the plurality of groups, at block 330, a set of shared data patterns for sharing among the collection devices in a group is determined based on the sets of data patterns associated with the collection devices in the group. Here, a group can include one or more collection devices. In the following, a process for determining a set of shared data patterns for a group of collection devices will be described with reference to Figure 6 Further details of determining shared data patterns will be described.

[0035] Figure 6 A block diagram of a process 600 for obtaining a set of shared data patterns according to an example implementation of the present disclosure is schematically shown. Although Figure 6 The process is schematically shown for a group of collection devices 210, similar processes can be performed for each of the plurality of groups. As Figure 6 As shown, the group of collection devices 210 includes collection device 112-1, collection device 112-2,.... Collection device 112-1 is associated with a set of data patterns 610 (including data patterns 612, 614,..., and 616), and collection device 112-2 is associated with a set of data patterns 620 (including data patterns 622, 624,..., and 626).

[0036] According to one example implementation of the present disclosure, the occurrence number of each data pattern in the set of data patterns in the data from each collection device in the set can be determined respectively. Specifically, the occurrence number of data patterns 612, 614, …, and 616 in the data from collection device 112-1 can be determined respectively, and the occurrence number of data patterns 622, 624, …, and 626 in the data from collection device 112-2 can be determined respectively. Further, based on the occurrence number of each data pattern, a set of shared data patterns 212 can be determined. For example, data patterns 612, 614, …, 616, 622, 624, …, and 626 can be sorted in descending order of occurrence number, and a specified number of data patterns with the highest ranking can be selected.

[0037] It will be appreciated that the more the number of shared data patterns stored at the edge computing device, the more duplicated data can be removed in the process of processing the data from collection devices 112. In addition, the storage space of edge computing device 114 is limited, which results in that too many shared data patterns cannot be stored at edge computing device 114. At this time, the number of duplicated data removed and the storage space occupied by shared data patterns should be balanced, and then the number of the set of shared data patterns is determined.

[0038] According to one example implementation of the present disclosure, the set of shared data patterns can also be determined based on the intersection of each set of data patterns. Specifically, the intersection between data patterns 612, 614, …, 616 and data patterns 622, 624, …, and 626 can be determined, so as to determine the data patterns that can be shared between collection device 112-1 and collection device 112-2.

[0039] According to one example implementation of the present disclosure, method 300 can be executed at any computing device in application environment 100. For example, method 300 can be executed at edge computing device 114 in edge network 110. At this time, the data from collection devices 112 can be processed internally in edge network 110, and then the shared data patterns are determined. In this way, the high bandwidth caused by transmitting the data to core network 120 can be avoided. For another example, method 300 can be executed at a server device in core network 120. It will be appreciated that the server device in core network 120 usually has rich computing resources, in this way, the data patterns can be processed with higher efficiency, and the shared data patterns that can be shared between each set of collection devices are determined.

[0040] According to one example implementation of the present disclosure, the set of shared data patterns determined based on method 300 can be distributed to edge computing devices 114 in edge network 110. In the following, referring to Figure 7A process of deploying a set of shared data patterns is described. Figure 7 A block diagram of deployment 700 of a set of shared data patterns according to an example implementation of the present disclosure is schematically shown. As shown, Figure 7 A set of shared data patterns 212 can be deployed to an edge computing device 114 in an edge network 110. Here, the edge computing device 114 is connected to a target collection device (e.g., collection device 112-1, 112-2, etc.) in a plurality of collection devices in a group.

[0041] Having described above the process of determining a set of shared data patterns for each group of collection devices, further, the edge computing device 114 can generate deduplicated data of target data from the target collection device based on the set of shared data patterns. Reference will be made to Figure 8 A process of removing duplicate data from data from collection devices 112 based on shared data patterns is described. Figure 8 A block diagram of a process 800 for generating deduplicated data according to an example implementation of the present disclosure is schematically shown. As shown at 8, data 210 from collection devices 112 can include a plurality of data segments 810, 812, …, and 814. Each data segment can be compared with a set of shared data patterns in order to determine whether each data segment can be replaced with an identifier of a shared data pattern.

[0042] If a given data segment matches one shared data pattern, the given data segment can be replaced with an identifier of the shared data pattern. If a given data segment does not match any shared data pattern, the given data segment remains unchanged. Figure 8 Deduplicated data 820 is shown, which can include an identifier 830 of a shared data pattern (as shown by arrow 840, the identifier 830 of the shared data pattern is used to replace data segment 810), data segment 812, …, and an identifier 834 of a shared data pattern (as shown by arrow 842, the identifier 834 of the shared data pattern is used to replace data segment 814).

[0043] With example implementations of the present disclosure, data segments 810 and 814, etc., in data 210 will be replaced with identifiers of certain shared data patterns in the set of shared data patterns. In this way, the amount of data of deduplicated data 820 will be less than the amount of data of data 210. This will reduce the overhead of storage resources and bandwidth resources involved in data storage and transmission.

[0044] According to one example implementation of the present disclosure, the edge computing device 114 can transmit the deduplicated data to a server device for processing the target data. Assuming that the server device 122 in the core network 120 is used to process data from a set of collection devices 210, the edge computing device 114 can transmit the deduplicated data from individual collection devices to the server device 122 that has removed the duplicated data. With the example implementation of the present disclosure, the bandwidth requirement between the edge network 110 and the core network 120 can be reduced, and the workload of data transmission can be reduced.

[0045] In the foregoing, reference has been made Figures 2 to 8 Examples of the method according to the present disclosure are described in detail, and implementations of the corresponding apparatus will be described hereinafter. According to an example implementation of the present disclosure, an apparatus for managing data patterns is provided. The apparatus comprises: an obtaining module configured to obtain a plurality of sets of data patterns respectively associated with a plurality of collection devices, a set of data patterns of the plurality of sets of data patterns representing a pattern of duplicated data in data from one collection device of the plurality of collection devices; a dividing module configured to divide the plurality of collection devices into a plurality of groups based on clustering of the plurality of sets of data patterns; and a determining module configured to determine a set of shared data patterns for sharing between the collection devices in a group of the plurality of groups based on the sets of data patterns associated with the collection devices in the group. According to an example implementation of the present disclosure, the apparatus further comprises modules for performing other steps in the method 300 described above.

[0046] Figure 9 A block diagram of an apparatus 900 for managing data patterns according to an example implementation of the present disclosure is schematically shown. As shown, the apparatus 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 902 or computer program instructions loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the apparatus 900 can also be stored in the RAM 903. The CPU 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0047] A plurality of components in the apparatus 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the apparatus 900 to exchange information / data with other apparatuses through a computer network, such as the Internet, and / or various telecommunication networks.

[0048] The various processes and processes described above, such as the method 300, can be performed by the processing unit 901. For example, in some implementations, the method 300 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some implementations, portions or all of the computer program can be loaded onto and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded onto the RAM 903 and executed by the CPU 901, one or more steps of the method 300 described above can be performed. Alternatively, in other implementations, the CPU 901 can also be configured in any other suitable manner to implement the above-described processes / methods.

[0049] According to an example implementation of the present disclosure, there is provided an electronic device comprising: at least one processor; a volatile memory; and a memory coupled with the at least one processor, the memory having stored therein instructions which, when executed by the at least one processor, cause the device to perform a method for managing data patterns. The method comprises: obtaining a plurality of sets of data patterns respectively associated with a plurality of collection devices, a set of data patterns of the plurality of sets of data patterns representing a pattern of repeated data in data from a collection device of the plurality of collection devices; dividing the plurality of collection devices into a plurality of groups based on clustering of the plurality of sets of data patterns; and determining a set of shared data patterns for sharing between the collection devices of a group of the plurality of groups based on the sets of data patterns respectively associated with the collection devices of the group.

[0050] According to an example implementation of the present disclosure, dividing the plurality of collection devices into the plurality of groups comprises: generating a plurality of pattern features of the plurality of sets of data patterns respectively based on the plurality of sets of data patterns; and dividing the plurality of collection devices into the plurality of groups based on the plurality of pattern features.

[0051] According to an example implementation of the present disclosure, generating the plurality of pattern features of the plurality of sets of data patterns respectively comprises: for a set of data patterns of the plurality of sets of data patterns, generating a pattern feature of the set of data patterns based on the data patterns of the set of data patterns and a number of occurrences of each data pattern in the data.

[0052] According to an example implementation of the present disclosure, dividing the plurality of collection devices into the plurality of groups comprises: converting the plurality of pattern features into a plurality of low-dimensional features respectively, the plurality of low-dimensional features having a lower dimension than the plurality of pattern features; and determining the plurality of groups based on clustering of the plurality of low-dimensional features.

[0053] According to an example implementation of the present disclosure, determining the set of shared data patterns comprises:

[0054] determining, for each group of data patterns, a number of occurrences of data patterns in the data from the respective collection devices in the group; and determining a set of shared data patterns based on the number of occurrences.

[0055] According to one example implementation of the disclosure, determining a set of shared data patterns includes determining a set of shared data patterns based on an intersection of the groups of data patterns.

[0056] According to one example implementation of the disclosure, obtaining a plurality of groups of data patterns includes obtaining a set of initial data patterns associated with the collection device; ranking the set of initial data patterns based on a number of occurrences of the set of initial data patterns in the data; and selecting a set of data patterns associated with the collection device based on the ranked set of initial data patterns.

[0057] According to one example implementation of the disclosure, the plurality of collection devices are located in an edge network in an application environment, and the method is performed at a computing device in the network.

[0058] According to one example implementation of the disclosure, the method further includes distributing the set of shared data patterns to an edge computing device in the edge network, the edge computing device being connected to a target collection device in the plurality of collection devices in the group; and instructing the edge computing device to generate deduplicated data for target data from the target collection device based on the set of shared data patterns.

[0059] According to one example implementation of the disclosure, the method further includes instructing the edge computing device to transmit the deduplicated data to a server device for processing the target data.

[0060] According to example implementations of the disclosure, a computer program product is provided that is tangibly stored on a non-transient computer readable medium and includes machine executable instructions for performing a method according to the disclosure.

[0061] According to example implementations of the disclosure, a computer readable medium is provided. The computer readable medium has stored thereon machine executable instructions that, when executed by at least one processor, cause the at least one processor to implement a method according to the disclosure.

[0062] The disclosure can be a method, apparatus, system, and / or computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for performing various aspects of the disclosure.

[0063] Computer readable storage media can be tangible storage media which can retain and store instructions for use by an instruction execution device. Computer readable storage media can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer readable storage media include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0064] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0065] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some implementations, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0066] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0067] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other data storage device. When the computer readable program instructions are loaded into the computer and other programmable data processing apparatus, a series of operational steps are implemented that provide processes such that the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0068] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0069] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0070] The above-described implementations of this disclosure are illustrative and not exhaustive, and are not limited to the disclosed implementations. Numerous modifications and adaptations will be apparent to those skilled in the art without departing from the scope and spirit of the disclosed implementations. The choice of words in this document is intended to best explain the principles of the implementations, practical application, or improvement over the technology in the art, or to enable other ordinary skilled in the art to understand the implementations disclosed herein.

Claims

1. A method for managing data schemas, comprising: Acquire multiple sets of data patterns associated with multiple acquisition devices located in an edge network of the application environment, and one set of data patterns represents a pattern of repetitive data from one of the acquisition devices. Multiple pattern features are generated based on the multiple sets of data patterns, wherein each pattern feature is generated for a corresponding set of data patterns in the multiple sets of data patterns, and wherein each pattern feature includes the number of occurrences of each individual data pattern in the corresponding set of data patterns. Based on the aforementioned multiple pattern characteristics, the multiple acquisition devices are divided into multiple groups; Based on the number of occurrences of each data pattern included in each group of data patterns associated with each acquisition device in each of the plurality of groups, a set of shared data patterns for sharing among the acquisition devices in the group is determined. The set of shared data patterns is distributed to edge computing devices in the edge network, wherein the edge computing devices are connected to a target acquisition device among the plurality of acquisition devices in the group; The edge computing device is instructed to generate deduplicated data of the target data from the target acquisition device based on the set of shared data patterns, wherein the amount of deduplicated data is smaller than the amount of the target data. The edge computing device is instructed to transmit the deduplicated data to the server device used to process the target data; as well as Therefore, the transmission of deduplicated data to the server device reduces the storage resource overhead involved in data storage by the server device.

2. The method according to claim 1, wherein dividing the plurality of acquisition devices into the plurality of groups comprises: The multiple pattern features are respectively transformed into multiple low-dimensional features, wherein the dimensions of the multiple low-dimensional features are lower than the dimensions of the multiple pattern features; as well as The multiple groups are determined by clustering based on the multiple low-dimensional features.

3. The method according to claim 1, wherein determining the set of shared data patterns further comprises: Based on the intersection of the various sets of data patterns, the set of shared data patterns is determined.

4. The method according to claim 1, wherein obtaining the multiple sets of data patterns includes: Acquire a set of initial data patterns associated with the acquisition device; Based on the number of times the initial data pattern appears in the data, the initial data pattern is sorted. as well as Based on the sorted initial set of data patterns, the set of data patterns associated with the acquisition device is selected.

5. The method of claim 1, wherein the server device is located in a core network separate from the edge network; and Therefore, the transmission of deduplicated data to the server device also reduces the bandwidth requirements between the edge network and the core network.

6. An electronic device, comprising: At least one processor; Volatile memory; as well as A memory coupled to the at least one processor, the memory having instructions stored therein, which, when executed by the at least one processor, cause the device to perform a method for managing a data pattern, the method comprising: Acquire multiple sets of data patterns associated with multiple acquisition devices located in an edge network of the application environment, and one set of data patterns represents a pattern of repetitive data from one of the acquisition devices. Multiple pattern features are generated based on the multiple sets of data patterns, wherein each pattern feature is generated for a corresponding set of data patterns in the multiple sets of data patterns, and wherein each pattern feature includes the number of occurrences of each individual data pattern in the corresponding set of data patterns. Based on the aforementioned multiple pattern characteristics, the multiple acquisition devices are divided into multiple groups; Based on the number of occurrences of each data pattern included in each group of data patterns associated with each acquisition device in each of the plurality of groups, a set of shared data patterns for sharing among the acquisition devices in the group is determined. The set of shared data patterns is distributed to edge computing devices in the edge network, wherein the edge computing devices are connected to a target acquisition device among the plurality of acquisition devices in the group; The edge computing device is instructed to generate deduplicated data of the target data from the target acquisition device based on the set of shared data patterns, wherein the amount of deduplicated data is smaller than the amount of the target data. The edge computing device is instructed to transmit the deduplicated data to a server device used to process the target data; and Therefore, the transmission of deduplicated data to the server device reduces the storage resource overhead involved in data storage by the server device.

7. The device according to claim 6, wherein dividing the plurality of acquisition devices into the plurality of groups comprises: The multiple pattern features are respectively transformed into multiple low-dimensional features, wherein the dimensions of the multiple low-dimensional features are lower than the dimensions of the multiple pattern features; as well as The multiple groups are determined by clustering based on the multiple low-dimensional features.

8. The device of claim 6, wherein determining the set of shared data patterns further includes: Based on the intersection of the various sets of data patterns, the set of shared data patterns is determined.

9. The device according to claim 6, wherein acquiring the plurality of data patterns includes: Acquire a set of initial data patterns associated with the acquisition device; Based on the number of times the initial data pattern appears in the data, the initial data pattern is sorted. as well as Based on the sorted initial set of data patterns, the set of data patterns associated with the acquisition device is selected.

10. The device according to claim 6, wherein The server device is located in a core network separate from the edge network; and Therefore, the transmission of deduplicated data to the server device also reduces the bandwidth requirements between the edge network and the core network.

11. A computer program product tangibly stored on a non-transient computer-readable medium and comprising machine-executable instructions for performing the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Data import method and device

    CN107463661A

  • Multi-domain convolutional neural network

    US20190042870A1