Cloud backup method and system for communication software

By collecting the device access logs and identifying duplicate path behaviors, a device behavior path mapping table is formed, the time fields are revised uniformly, and multi-level inode nodes are built, which solves the problems of inaccurate incremental identification and low backup efficiency in the existing technology, and improves the accuracy and reliability of data backup.

CN120386669AActive Publication Date: 2025-07-29SHENZHEN BESTONE TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510876672.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In the context of parallel use of multiple devices, the data files are extracted by accessing local storage directories and using timestamp comparison to incremental recognition, resulting in inaccurate incremental recognition, missing data or repeated backups, loose backup data organization, inefficient retrieval, and failure to effectively associate device behavior paths and data item characteristics, resulting in slow backup recovery and reduced reliability.

Method used

Collect device access logs, extract node jump sequences, identify duplicate path combinations, filter node data combinations through the order consistency of device paths and data labels, form a continuous mapping backup scheduling chain, uniformly revise the device record time fields, arrange structured index path sets in combination with field access order, and build multi-level inode nodes to ensure the orderliness of data writing and retrieval efficiency.

Benefits of technology

It improves the accuracy and retrieval efficiency of data organization, avoids data redundancy and backup conflicts, enhances the reliability and query efficiency of data recovery, and ensures the overall availability and reliability of the data backup system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386669A_ABST
    Figure CN120386669A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data management, in particular to a cloud backup method and system for communication software, and the method comprises the following steps: collecting differentiated equipment access logs based on a cloud backup node connected by a user, extracting a node jump sequence to recognize repeated combinations, comparing data labels according to a path sequence, screening consistent combinations, and classifying the consistent combinations into a scheduling channel; and extracting a communication record revision time field, arranging a field structure in combination with an access sequence, and constructing a multi-level index node to generate a backup index tree configuration result. According to the method, the repeated path behavior is recognized by collecting the device access log and extracting the node jump sequence, the accuracy of data organization is effectively improved, session data time synchronization is achieved, the structured index path set is arranged in combination with the field access sequence, and the multi-level index nodes are constructed according to the path positions; the data writing orderliness and retrieval efficiency are ensured, data redundancy and backup conflicts are avoided, and the data recovery reliability and query efficiency are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data management, and in particular to a cloud backup method and system for a communication software. Background Art

[0002] The technical field of data management involves core matters such as the collection, storage, organization, maintenance, backup, recovery, and access control of various types of data. It covers the full life cycle management technology of data, supports the structured and unstructured processing of data, access logic control, data consistency guarantee, disaster tolerance mechanism construction, data synchronization and version management, cloud storage scheduling, etc., and supports the security, integrity, and availability of data in different systems or platforms.

[0003] Among them, the cloud backup method of traditional communication software refers to a solution for remotely storing user communication data to ensure data security. The technical matters targeted are how to effectively synchronize user data such as local message records, pictures, voices, videos, etc. to a remote server for periodic incremental backup in a mobile communication software environment. After extracting data files by authorizing access to the local storage directory of the device, they are transmitted to the cloud server through the HTTPS protocol. At the same time, incremental recognition is performed by comparing the timestamps recorded in the database, and updated data is retrieved. The object storage service is used for classified storage and index management.

[0004] The prior art extracts data files by accessing the local storage directory and uses timestamp comparison for incremental recognition. It relies on a single time field to determine data updates. In the scenario of parallel use of multiple devices, it is easy to cause inaccurate incremental recognition due to inconsistent time fields, resulting in problems such as data omission or duplicate backup. Moreover, the traditional method fails to effectively associate the device behavior path with the data item characteristics, resulting in loose organization of backup data and low retrieval efficiency. In long-term backup and recovery, it is easy to cause data chaos or slow recovery, reducing the overall availability and reliability of the data backup system. Summary of the Invention

[0005] The purpose of the present invention is to solve the deficiencies existing in the prior art, and to propose a cloud backup method and system for a communication software.

[0006] To achieve the above purpose, the present invention adopts the following technical solution. A cloud backup method for a communication software includes the following steps: S1: Based on the cloud backup nodes connected by users in the communication software, collect the access log data of the nodes on different devices, extract the node jump sequence in the continuous access behavior according to the device number, identify duplicate node combinations, and obtain the device behavior path mapping table; S2: Call the device paths in the device behavior path mapping table, sequentially detect the client tags bound in the associated data items according to the path order, compare the tags of the data items with the nodes in the path, filter out the node data combinations whose behavior types are consistent with the path order, and output the cloud backup scheduling node chain group; S3: Call the communication records in each channel of the cloud backup scheduling node chain group, extract the device number, session identifier, and message sending time, classify them by device, locate the start node time field of the session, perform time comparison, and uniformly revise the time field of the device record to obtain the unified session data time table; S4: Call the field information of each data item in the unified session data time table, extract the fields of sending time, message length, file format, and session number, combine the access order of the corresponding fields, and arrange the field structure according to the access sequence to form a field order index path set.

[0007] As a further solution of the present invention, the device behavior path mapping table includes a path number, a node jump mode, and a device behavior tag. The cloud backup scheduling node chain group includes a channel identifier, a node binding mapping, and a path sequence number. The unified session data time table includes a message time mapping table, a device unified timestamp, and a synchronization status identifier. The field order index path set includes a field order list, a field position index, and an access structure tag.

[0008] As a further solution of the present invention, the obtaining steps of the device behavior path mapping table are specifically as follows: S111: Based on the cloud backup nodes connected by users in the communication software, collect the access log data on different devices, extract the device number and the complete sequence of corresponding access behaviors, arrange the log records in the order of access, and obtain the device access behavior sequence group; S112: Call the device access behavior sequence group, for the continuous access behavior sequences under the same device number, extract the node number field in the access records, compare the node jump paragraphs, identify the jump segments with the same node combination, and perform number mapping on the repeated jump segments to obtain the node jump combination set; S113: According to the node jump combination set, accumulate the occurrence frequencies of the repeated combinations in the node jump paths of each device, combine the node jump path length and the position of the jump combination in the path, calculate the mapping intensity value of the device in the node combination path, summarize the mapping intensity values by device number, analyze the node behavior relationship of the device, and obtain the device behavior path mapping table.

[0009] As a further solution of the present invention, the obtaining steps of the cloud backup scheduling node chain group are specifically as follows: S211: Invoke the device paths in the device behavior path mapping table, sequentially detect the client tags bound in the associated data items according to the path order, compare the tags of the data items with the nodes in the path, screen the node data combinations according to the consistency between the data item tags and the node tags, and obtain the node data combinations with consistent tags; S212: Based on the node data combinations with consistent tags, screen the node data combinations whose behavior types are consistent with the path order, calculate the behavior order matching degree value, screen the node data combinations that meet the requirements according to the behavior order matching degree value, and generate the sequentially matching node data combinations; S213: Based on the sequentially matching node data combinations, combine the continuous mapping relationship between the nodes and the data items, classify the node chain groups that meet the mapping relationship into the scheduling channels, and output the cloud backup scheduling node chain groups.

[0010] As a further solution of the present invention, the step of obtaining the session data unified time table is specifically as follows: S311: Based on the communication records of each channel in the cloud backup scheduling node chain group, extract the device number, session identifier and message sending time, classify the different device numbers, and aggregate the communication records with the same device number to obtain the device communication record set; S312: Invoke the device communication record set, for each session identifier in the set, locate the start node time field of the session, extract the start node time field data, compare the start node times of different device numbers, calculate the unified time difference metric value between devices, uniformly revise the device record time fields, and combine the revised time field sets to integrate the time records of each session identifier to generate the session data unified time table.

[0011] As a further solution of the present invention, the step of obtaining the field order index path set is specifically as follows: S411: Invoke the field information of multiple pieces of data in the session data unified time table, extract the fields of the sending time, message length, file format and session number recorded in each piece of data, and for each piece of data, sequentially extract the corresponding fields according to the field identification position to obtain the field combination set; S412: Use the content of the fields of the sending time, message length, file format and session number in the field combination set, combine the field access order information recorded in the client access log data, compare it with the field access order number under the same number in the access log, and combine the real-time arrangement sequence of each group of data fields to obtain the field order index path set.

[0012] As a further solution of the present invention, the method further includes the step: S5: Invoke the path positions corresponding to the fields in the field sequence index path set, successively construct multi-level index nodes, complete the configuration of upper-layer, middle-layer, and lower-layer nodes in the path order, bind each data write request to the corresponding structural path, and generate a backup index tree configuration result; The backup index tree configuration result includes an index node structure, a data binding relationship, and a structure level identifier.

[0013] As a further solution of the present invention, the steps for obtaining the backup index tree configuration result are specifically as follows: S511: Invoke the path positions corresponding to the fields in the field sequence index path set, and according to the field arrangement order, successively locate the position information of the fields in the path set, and combine the position information with the field names to generate a field path index sequence; S512: Based on the field path index sequence, successively construct multi-level index nodes, aggregate the fields with the same path positions at the same level into the same node, respectively set the structural relationships of the upper layer, middle layer, and lower layer, and generate an index level node connection structure; S513: Invoke the index level node connection structure, perform path recognition processing on each data write request, bind the data content to the corresponding node positions according to the field mapping relationship, and sequentially nest and construct a tree structure according to the structure level to generate a backup index tree configuration result.

[0014] The cloud backup system of the communication software is used to execute the above-mentioned cloud backup method of the communication software, and the system includes: The node access and collection module extracts the device number, node address, and access time fields based on the cloud backup nodes connected by users in the communication software, groups them according to the device number, arranges the node order according to the access time, judges the continuous access node jump, and generates a device behavior path mapping table; The behavior path mapping module uses the device behavior path mapping table to identify the node combinations in the jump sequence, counts the occurrence frequency of the node combinations, and combines with the device number to obtain a cloud backup scheduling node chain group; The node label screening module invokes the cloud backup scheduling node chain group, successively extracts the client labels bound to the associated data items based on the path order, and generates a session data unified time table by comparing the node labels with the data item labels one by one; The session time unification module invokes the communication records of the channels in the session data unified time table, extracts the device number, session identifier, and message sending time, classifies the data according to the device number, locates the sending time of each session, and generates a field sequence index path set; The backup index module calls the set of field order index paths, extracts the sending time, message length, file format, and session number, sorts the fields based on the sending time, extracts the field path positions, constructs index nodes in the path order, binds the data records, and obtains the backup index tree configuration result.

[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In the present invention, by collecting device access logs and extracting node jump sequences, the recognition of repeated path behaviors is achieved. Based on the sequential consistency of the device paths and data tags, the node data combinations are screened to form a continuously mapped backup scheduling chain, effectively improving the accuracy of data organization. By comparing the time of communication records, the device record time fields are uniformly revised to achieve session data time synchronization. Combining the structured index path set arranged according to the field access order, multi-level index nodes are constructed according to the path positions to ensure the orderliness of data writing and retrieval efficiency. The overall process improves the structural integrity of data backup and the consistency of access timing, avoids data redundancy and backup conflicts, and enhances the reliability and query efficiency of data recovery. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a schematic diagram of the working process of the present invention; Figure 2 is a flowchart for obtaining the device behavior path mapping table in the present invention; Figure 3 is a flowchart for obtaining the cloud backup scheduling node chain group in the present invention; Figure 4 is a flowchart for obtaining the session data unified time table in the present invention; Figure 5 is a flowchart for obtaining the set of field order index paths in the present invention; Figure 6 is a flowchart for obtaining the backup index tree configuration result in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0018] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality of" is two or more unless otherwise specifically defined.

[0019] Please refer to Figure 1 , the present invention provides a technical solution, a cloud backup method for a communication software, including the following steps: S1: Based on the cloud backup nodes connected by users in the communication software, collect the access log data of the nodes on different devices, extract the node jump sequence in the continuous access behavior according to the device number, identify the repeated node combinations, and obtain the device behavior path mapping table; S2: Call the device paths in the device behavior path mapping table, sequentially detect the client tags bound in the associated data items according to the path order, compare the tags of the data items with the nodes in the path, screen out the node data combinations with the behavior type consistent with the path order, classify the combinations into the scheduling channels, and combine the continuous mapping relationship between the nodes and the data items in the scheduling channels to output the cloud backup scheduling node chain group; S3: Call the communication records in each channel of the cloud backup scheduling node chain group, extract the device number, session identifier and message sending time, classify by device and then locate the start node time field of the session, perform time comparison and then uniformly revise the time field of the device record to obtain the unified session data time table; S4: Call the field information of each data in the unified session data time table, extract the fields of sending time, message length, file format and session number, combine the access order of the corresponding fields in the client access log data, and arrange the field structure according to the access sequence to form a field order index path set; S5: Call the path positions corresponding to the fields in the field order index path set, sequentially construct multi-level index nodes, complete the configuration of the upper, middle and lower layer nodes according to the path order, bind each data write request to the corresponding structure path, and generate the backup index tree configuration result; The device behavior path mapping table includes the path number, node jump mode, and device behavior label. The cloud backup scheduling node chain group includes the channel identifier, node binding mapping, and path sequence number. The session data unified time table includes the message time mapping table, device unified timestamp, and synchronization status identifier. The field sequence index path set includes the field sequence list, field position index, and access structure label. The backup index tree configuration result includes the index node structure, data binding relationship, and structure hierarchy identifier.

[0020] See also Figure 2 , the specific steps for obtaining the device behavior path mapping table are: S111: Based on the cloud backup node connected by the user in the communication software, access log data on different devices is collected, the complete sequence of device numbers and corresponding access behaviors is extracted, and the log records are arranged in order of access to obtain a device access behavior sequence group; The access log data on differentiated devices is collected, and all bound terminal devices are located through user identification. A unique number index is established for each device in turn, and devices A, B, and C are marked as ID001, ID002, and ID003 respectively. The log entries corresponding to each number are then extracted one by one from the cloud node database. The extracted fields include: access timestamp, target node number, access type, and authentication status. The timestamp is accurate to milliseconds, the access type is GET / POST, and the authentication status is recorded using a Boolean variable. The log entries are classified by device number and sorted in ascending order according to the timestamp field to obtain the continuous access records of each device in the order of actual operation. Regarding the behavior sequence, the access records generated by device ID001 within a certain period of time are as follows: node N001 accessed at 08:12:23.345, node N002 accessed at 08:12:24.182, and node N003 accessed at 08:12:27.981. The three access behaviors constitute the access behavior sequence {N001→N002→N003} in chronological order. To prevent data anomalies from interfering with the behavior path, entries with an authentication status of False are removed from the access behavior, and only valid access behaviors after successful user authentication are retained. After removing invalid records, the node integrity of each access sequence is reconfirmed, and the device number is used as the index item to obtain the device access behavior sequence group.

[0021] S112: Calling the device access behavior sequence group, extracting the node number field in the access record for the continuous access behavior sequence under the same device number, comparing the node jump segments, identifying jump segments with the same node combination, and performing number mapping on the repeated jump segments to obtain a node jump combination set; For each group of device access sequences in the set, extract the node number fields in the access behavior records, form a jump path list in the access order, compare the node combinations of each jump path in different devices. Set device ID001 to have a path sequence {N001→N002→N003→N004}, ID002 to be {N001→N002→N003→N005}. There is a repeated node combination {N001→N002→N003} between them. Then this combination can be marked as a repeated path segment. Judge whether the repeated combination has a continuous node relationship with a path segment length greater than or equal to 3. If the condition is met, extract this segment combination to establish a standard path index table, map and label the device paths, compare and number the repeated segments that meet this feature in the path. For example, {N001→N002→N003} is assigned a path segment label R01. At the same time, record the start position and end position of this path segment in the original access sequence. Then the access path can be transformed into a logical jump path sequence composed of a mixture of ordinary nodes and combination labels. In the path segment determination, set the validity threshold of the repeated combination to the occurrence frequency ≥2. If the number of occurrences of a certain node combination in different device paths is lower than this threshold, it is not regarded as a repeated path segment. Set the combination {N002→N004} to appear only 1 time in ID001, then it is excluded from the valid node combinations. Construct the combination numbers and structures of all combinations that meet the requirements of path segment length and repetition frequency in a list form to obtain a node jump combination set.

[0022] S113: According to the node jump combination set, accumulate the occurrence frequencies of the repeated combinations in the node jump paths of each device. Combining the node jump path length and the position of the jump combination in the path, use the formula: ; Calculate the mapping intensity value of the device in the node combination path, summarize the mapping intensity values by device number respectively, analyze the node behavior relationship of the device, and obtain a device behavior path mapping table; Among them, represents the node behavior mapping intensity value of device , represents the occurrence frequency of the th node combination in device , represents the path length of the th node combination in device , represents the jump position serial number in the path of the th node combination in device , represents the jump density value of the th node combination in device , is the number of node combinations in the device; Formula calculation logic: The formula is used to calculate the node behavior mapping intensity value of the device and its core logic consists of two parts. The first part is the product of the frequency term and the structure term, that is , which reflects the product of the number of times a node combination appears in a path and the structure length, and then divided by the relative position of the combination in the path The square root is used to weaken the influence of subsequent combinations in the path and form an attenuation mechanism; the second part is the average value of the jump density value of the node combination in the device , indicating the density of jumps in the path. The calculation results of the two parts are subtracted and the absolute value is taken to ensure that regardless of whether the frequency structure is high or the jump density is low, the mapping abnormality degree can be measured. Overall, the formula forms a composite scoring mechanism by integrating four types of parameters: the frequency of structure appearance, length, position, and jump tightness, and quantitatively evaluates the structural stability and access consistency of path combinations, which is applicable to path structure comparison and classification of multiple devices under the same logical service; The node behavior mapping intensity value is a composite index that measures the structural characteristics and jump distribution of repeated node combinations in the device access path. By calculating the frequency of appearance, path length, path position, and jump density of the node combination, the structural characteristics of its behavior pattern are extracted. The higher the value, the more dense and clearly structured the repeated jump combination pattern of the device in the access path, which is convenient for path normalization mapping and abnormal behavior recognition; Calculate the reciprocal mean of the number of jumps between adjacent combination segments in the same device and set it as the jump density value , if the total number of jumps in the path is 12 and the distances between repeated combination segments are 2, 2, 4, 4 respectively, then the jump density is , substitute the above four parameters into the formula: Taking device ID001 as an example, it has 3 repeated path combinations R01, R03, R05, and the parameter settings are as follows: Table 1 Device parameter table As shown in Table 1, the substitution calculation process is as follows: Calculate each product and divide by the square root: R01: ; R03: ; R05: ; Calculate the sum and subtract the average jump density: Transfer item sum: Average density term: Jump intensity value: ; Take the obtained Recorded in the node behavior mapping values of ID001, construct the node behavior atlas for each device number in sequence, and generate the device behavior path mapping table; Table 2 gives the example parameters used in the above calculation process: Table 2 Device Node Combination Statistical Parameter Table As shown in Table 2, different node combination items have clear structured records in the device path, which is convenient for subsequent unified structure scoring and difference quantification; The result shows that through the calculation of the node behavior mapping intensity value, a unique identifier and structure characterization parameter can be provided for the device behavior path mapping table, forming a device path structure description result with comparability and decidability.

[0023] Please refer to Figure 3 , the steps for obtaining the cloud backup scheduling node chain group are specifically as follows: S211: Call the device path in the device behavior path mapping table, sequentially detect the client tags bound in the associated data items according to the path order, compare the data items with the nodes in the path for tag comparison, and screen the node data combinations according to the consistency of the data item tags and node tags to obtain the node data combinations with consistent tags; Based on the path information extracted from the device behavior path mapping table, obtain the device identification information of the nodes in the path. The device identification information includes the device ID, node sequence index, and bound client tag. Obtain the associated data items. The data items record the bound client tag and data item encoding. Use a traversal method to sequentially detect the client tags bound to each data item according to the path order. For each node, call its device identification information and compare it with the client tag in the data item one by one. Use a matching determination. If the client tag ID of the node is the same as the client tag ID in the data item, record the corresponding relationship between the node and the data item, and combine the qualified nodes and data items into a preliminary node data combination. In the example, set the client tag IDs bound to nodes A, B, and C in the path to T1, T2, and T3 respectively, and the client tag IDs bound to data items D1, D2, and D3 are T1, T3, and T2 in sequence. Then through comparison, node A corresponds to data item D1, node B corresponds to data item D3, and node C corresponds to data item D2, forming three groups of node data combinations respectively. During the detection process, process them in ascending order of the node sequence index number to avoid sequence chaos, and during the comparison process, use a strict equality determination. The client tag ID needs to be exactly the same and does not support fuzzy matching. If the client tag ID uses a 16-bit UUID encoding method, then compare bit by bit. The preliminary node data combination can be listed in the following table: Table 3 Node and Data Item Tag Comparison Table As shown in Table 3, by comparing the client label IDs of nodes and data items, the node data combinations with consistent labels were determined, and the node data combinations with consistent labels were obtained.

[0024] S212: Based on the node data combinations with consistent labels, filter out the node data combinations with consistent behavior types and path orders, and use the formula: ; Calculate the behavior order matching degree value, filter out the node data combinations that meet the requirements according to the behavior order matching degree value, and generate the order matching node data combinations; Among them, is the behavior order matching degree value, represents the behavior type code bound to the node, represents the sequential index value of the node in the path, represents the timestamp of the data item, represents the client label ID, represents the node label ID, represents the number of nodes in the path, represents the number of node combinations with consistent labels; The calculation logic of the formula is as follows: Sum the behavior type code bound to the node, the sequential index value , and the data item timestamp respectively to form a unified sum of behavior time characteristics, reflecting the overall distribution characteristics of nodes in the path in terms of behavior and time; Perform a difference operation on the client label set of the data item and the client label set of the node respectively. The difference reflects the consistency degree between the path nodes and the data item in the label dimension. The smaller the difference, the higher the label matching degree; Multiply the sum of the behavior time characteristics by the label difference result, and combine the sum of the number of path nodes and the number of node combinations with consistent labels for normalization processing. After normalization, take the absolute value and square root to output the behavior order matching degree value , eliminating the influence of a single data dimension on the overall matching degree, ensuring the unity of the dimensions of each participating quantity, and the calculation result reflects the overall matching degree of the node path in terms of behavior characteristics and data consistency; The behavior order matching degree value is used to measure the consistency degree of the path order and time characteristics of the device node based on the consistency of the behavior type and the data item label. This value synthesizes the node behavior coding, sequential index, timestamp, and label matching degree, and quantifies the overall matching relationship of the node combination through unified normalization operations. The closer the behavior order matching degree value is to 0, the higher the consistency of the node data in terms of behavior and order, and vice versa, indicating a low matching degree; Extract the behavior type codes and sequential index values bound to each node. At the same time, extract the timestamps of the data items and the client label IDs, and then construct data sets respectively. and and and and and and , where represents the set of node behavior type codes, represents the set of sequential indices of path nodes, represents the set of timestamps of data items, represents the set of client label IDs, represents the set of node label IDs, represents the number of nodes within the path, represents the number of node combinations with consistent labels; Set the behavior type codes to 101, 102, 103 respectively, the sequential indices to 1, 2, 3, the timestamps to 1620000001, 1620000020, 1620000030, the client label IDs to be equal to the node label IDs respectively, with the values being 1001, 1002, 1003, the number of path nodes , and the number of combinations with consistent labels , then the specific calculation process is as follows: ; ; ; ; In the formula, represents the sum of the node behavior type code, sequential index, and data item timestamp, represents the sum of the differences between the set of client label IDs and the set of node label IDs, represents the sum of the total number of nodes and data items. From the above calculations, the behavior sequence matching degree value is obtained. Since the value is zero and does not meet the subsequent screening requirements, interference value adjustment needs to be introduced. It is set that in actual applications, there is an error adjustment term for the behavior coding. After resampling, the coding is 101, 103, 105, and recalculate after the update: ; ; After the interference item is adjusted to zero, a reference value needs to be set for secondary screening. The reference value is set to 0.1, which is taken from the standard deviation of the average device operation delay. Through actual measurement of multiple groups of samples, the standard deviation of the device delay is about 0.08 to 0.12 seconds. Therefore, the reference value is set within the interval. Screening is carried out according to the reference value. If , then this combination is excluded. Through screening, an ordered matching node data combination is obtained.

[0025] S213: Based on the ordered matching node data combination, combined with the continuous mapping relationship between nodes and data items, the node chain groups that meet the mapping relationship are classified into the scheduling channel, and the cloud backup scheduling node chain group is output; Call the mapping relationship between nodes and data items to extract continuous node chain groups. A continuous node chain group means that the node IDs are continuously indexed in order, and the time difference between the corresponding data item timestamps does not exceed the set time limit. The set time limit value is 5 seconds, and the upper limit is set as 10 times the device acquisition cycle of 0.5 seconds. Set the node order and timestamp respectively as: Table 4 Node Chain Group Timestamp Table As shown in Table 4, the time difference between node B and node A is 4 seconds, which is less than 5 seconds, and the time difference between node C and node B is 5 seconds, which is equal to the time limit. Both meet the continuity requirements. The nodes A, B, and C are combined into a chain group in order, classified into the scheduling channel, and a cloud backup scheduling node chain group is generated.

[0026] Please refer to Figure 4 , the specific steps for obtaining the unified session data schedule are as follows: S311: Based on the communication records of each channel in the cloud backup scheduling node chain group, extract the device number, session identifier, and message sending time, classify the different device numbers, and aggregate the communication records with the same device number to obtain the device communication record set; Collect communication record data, extract the corresponding device number, session identifier, and message sending time fields one by one. For the device number field in the communication record data, perform data classification operations, that is, group the communication records with the same device number according to the device number and classify them into a set of device communication records. Taking the device number A123 as an example, after classification, the corresponding set contains several communication records, and the message sending times of the records are 2025-05-01 12:00:00, 2025-05-01 12:01:30, 2025-05-01 12:03:00, etc. By sequentially scanning the device number fields in each communication record and aggregating the records with the same device number into a set, in the case of a large amount of data, a doubly linked list can be used to store the set of device communication records to improve the subsequent traversal efficiency. For the device number A124, the sending times in the obtained set after classification are 2025-05-01 12:00:10, 2025-05-01 12:01:40, 2025-05-01 12:03:20, etc. After classification, the sets of device communication records are independently grouped and stored according to the device number, which is convenient for subsequent positioning and processing of the session identifier, and the set of device communication records is obtained.

[0027] S312: Invoke the set of device communication records. For each session identifier in the set, locate the start node time field of the session, extract the data of the start node time field, and compare the start node times under different device numbers. Use the formula: ; Calculate the unified time difference metric value between devices, uniformly revise the device record time fields, and combine the revised time field sets to integrate the time records of each session identifier to generate a unified time schedule for session data; Among them, represents the unified time difference metric value between device and device , represents the value of the start node time field of device , represents the value of the start node time field of device ; The calculation logic of the formula: Based on the values of the start node time fields of device and device , calculate the time difference between the two, and ensure that the time difference is non-negative through the absolute value operator to avoid the influence of the time sequence on the result stability. Take the square root after calculating the product of the two time fields, and use Reflect the closeness of the magnitude of the time fields of two devices. Eliminate the non-linear deviation introduced by the magnitude difference through square root processing, and then perform an addition operation on the absolute value difference and the square root value, so that the time difference and the magnitude of the time product are measured on the same scale. Use as the denominator for normalization processing, scale the total amount of the quantized time difference, and make the measurement value fall within a relatively stable range. This process comprehensively considers the differences and overall magnitudes of the time fields of the two devices, ensuring that the unified time difference measurement value has both sensitivity and robustness, and is applicable to scenarios where there are small offsets but the same magnitude in the time records of different devices, further improving the consistency and accuracy of time revision; The unified time difference measurement value is a standardized value used to measure the relationship between the starting node time difference and the overall magnitude of two devices in the same session. This measurement value combines the absolute value of the time difference and the square root of the time product, and through normalization processing, makes the result reflect the offset and consistency level of the time fields of the two devices. The closer the measurement value is to 0, the closer the starting times of the two devices are; the closer it is to 1, the relatively larger the time deviation is. It is suitable as a judgment basis for unified revision processing; Traverse the records in the set, group them according to the session identifier, and select the record with the earliest message sending time field in each group as the starting node time field of the session. Taking session C1 of device number A123 as an example, the starting time field is 2025-05-01 12:00:00, and the starting time field of session C1 of device number A124 is 2025-05-01 12:00:10. After extracting the starting node time field data of each device as above, perform pairwise time field comparison operations; Set the starting node time of device A123 to 43200 seconds, and the starting node time of device A124 to 43210 seconds. Substitute into the formula for calculation as follows: ; ; ; ; ; Thus, the unified time difference measurement value between devices is obtained as 0.4999. Further, set the unified revision threshold for the time difference to 0.5. If the value is lower than this threshold, then perform unified revision processing. The revision method is to take the average value of the starting node times of each device, that is: ; Revise the start time of each device to 43205 seconds. After revision, the times are 12:00:05 respectively. Based on the revised time field set, integrate the time records of each session identifier to generate a time record table in a unified format, forming a unified time schedule for session data.

[0028] Please refer to Figure 5 , and the specific steps for obtaining the field order index path set are as follows: S411: Invoke the field information of multiple pieces of data in the unified time schedule for session data, extract the fields of send time, message length, file format, and session number recorded in each piece of data. For each piece of data, extract the corresponding fields in sequence according to the field identifier position to obtain a field combination set; Perform segmented extraction on each piece of data in the time schedule. The invocation operation should be completed based on the field identifier structure, which contains the field name, field type, and the position value of the field offset. By reading the structure definition, the actual start offset position and length value of the four fields of send time, message length, file format, and session number in each piece of data can be located, and the content of each field is parsed and extracted in sequence. Among them, the send time can be recorded in the Unix timestamp format. Set the Unix timestamp 1617273600 of 4 bytes to be recorded at the position with the starting offset value of 0 in a certain piece of data, and the corresponding Beijing time is 0:00 on April 1, 2021; the message length field follows immediately. If it is set as a 2-byte unsigned integer, it can be parsed as a certain message containing a byte value such as 128 bytes; the file format field can be represented by fixed characters, such as ".txt" or ".doc", and is stored in ASCII code in the original record. The corresponding position offset value is 6, and the length is set to 4 bytes. After reading, it is converted into a character form; the session number field is set as a non-repeating integer identifier, such as session number 10086, which is set as a 4-byte integer field starting at the offset value of 10; after the above fields are extracted in sequence, the four fields are concatenated in the order of "send time - message length - file format - session number" according to the data structure order to form a field group. Repeat the extraction process for each record in the unified time schedule, and set the following field combination groups to be formed for three consecutive pieces of data: {1617273600, 128, ".txt", 10086}, {1617277200, 256, ".pdf", 10087}, {1617280800, 64, ".doc", 10088}, and summarize them to form a field combination set.

[0029] S412: Use the content of the send time, message length, file format, and session number fields in the field combination set, combine the field access order information recorded in the client access log data, compare it with the field access order number under the same number in the access log, and combine the real-time arrangement sequence of each group of data fields to obtain the field order index path set; Read the values of the sending time, message length, file format, and session number in each group one by one according to the structure of the field combination set, and match the session number with the corresponding record in the client access log data. By parsing the field access sequence number field in the client access log, identify the order in which the field corresponding to the session number is accessed on the client. Set the field access order of the record with the session number 10086 in the client access log to {3, 1, 4, 2}, which means that the field in the 3rd position in the original field order (i.e., the file format) is accessed first, followed by the field in the 1st position (i.e., the sending time), the 4th position (session number), and the 2nd position (message length). It is necessary to reorder the field sequence in the field combination set and arrange the structure of the current field combination group according to this order. That is, reorder the original field group {1617273600, 128, ".txt", 10086} to {".txt", 1617273600, 10086, 128}; repeat this matching and reordering operation, and perform a log access order parsing and field reordering operation on each field group in the field combination set to construct a new structure set composed of the reordered field sequences; it is necessary to perform structure encoding on the field order in each reordered structure. Use the first letter of the field type to represent the field type. Set S to represent the sending time, L to represent the message length, F to represent the file format, and C to represent the session number. Encode the aforementioned reordered order as "F-S-C-L" to construct a field order index path set.

[0030] Please refer to Figure 6 , and the specific steps for obtaining the backup index tree configuration result are as follows: S511: Call the path positions corresponding to the fields in the field order index path set. According to the field arrangement order, locate the position information of the fields in the path set in turn, and combine the position information with the field names to generate the field path index order; Extract the fields in the data structure and their arrangement order information in the structure path set. When processing the data records of each customer in the enterprise customer information setting, the fields include customer number, contact information, order status, etc., which exist in the form of nested objects in the original structure. Set the fields "customer.id", "customer.contact.phone", "customer.orders.orderID". At this time, the field paths included in the path set need to be extracted according to their hierarchy and order. Through the preset path extraction, read the position information of the fields in the structure in turn, clarify the nested hierarchy where each field is located and the position number at that level, and bind and combine the obtained path positions with the field names to form a field path sequence with position information. Set the field customer.contact.phone to be at path level 2 and sequence 3, which is represented as (2, 3, customer.contact.phone). Rearrange the above structure order pairs according to the initial definition order of the fields, and output the field path set arranged in order. During the processing, if it is found that a certain field has multiple levels or repeated path names in the path, its path needs to be extended and marked, and the unique position is indicated by appending a hierarchy prefix to avoid field index conflicts. This execution process outputs a set of structured field path index order data, which serves as the basis for subsequent node construction and generates the field path index order.

[0031] S512: Based on the field path index order, construct multi-level index nodes in sequence, aggregate the fields with the same path positions at the same level into the same node, respectively set the structural relationships of the upper layer, middle layer and bottom layer, and generate an index level node connection structure; It is necessary to classify the levels to which the fields in the path belong first, group the fields with the same path levels into one group, and set the field path order in financial transactions to include fields such as account.id, account.details.balance, account.details.limit, transaction.id, transaction.amount. Among the fields, "account.id" is a first-level field, "account.details.balance" and "account.details.limit" are second-level fields, and the "transaction"-related fields are another level of fields. Then, in the construction of the index structure, fields at the same level need to be aggregated to form nodes, and the superior-subordinate node relationship needs to be established according to the field path index order. By marking the path prefix as the connection clue, a multi-level structure relationship with "account" at the upper layer, "account.details" in the middle layer, and "balance" or "limit" at the bottom layer is formed. The nodes are bound through the path prefix and the index order. For example, if there are two fields, "account.details.limit" and "account.details.threshold", both belonging to the middle-layer node "account.details", they should be set as its subordinate bottom-layer nodes. During the process, each node should be marked with the number of fields it contains, the path name, and its index position in the data to form a clearly structured node chain structure. After the nodes are set up, each node is assigned a path recognition identifier and its subordinate relationship chain is recorded for path positioning and structure binding during subsequent nested data writing. After this execution process is completed, a clearly callable set of node structures is output to generate an index-level node connection structure.

[0032] S513: Call the index-level node connection structure, perform path recognition processing on each data write request, bind the data content to the corresponding node positions according to the field mapping relationship, and sequentially nest and construct a tree structure according to the structure level to generate a backup index tree configuration result; First, extract the field values of each record from the actual business data. Taking medical images as an example, the fields include patient.id, exam.date, exam.findings.description, etc. After receiving a data write request, according to the previously constructed index node structure, identify the field paths one by one for each field. According to the field path mapping relationship, extract the field values from the record. For example, when the field exam.findings.description is identified, the corresponding structure path is "description" under "findings" under the node "exam". Read the field value such as "The nodule image is clearly visible" from the data according to this structure path and bind it to the structure path node. After the field path binding is completed, construct the corresponding structure body of the field layer by layer according to the constructed multi-level structure relationship to form a complete tree structure. During the nesting process, if no field value is obtained under a certain path node, retain the null value node structure to maintain the structural consistency. After the entire tree structure is constructed, store it in the form of a configuration file or a structured document and perform index registration for data retrieval. The fields in the written structure all have the binding relationship between the path label and the content, which is convenient for quick positioning and reading in subsequent operations. After this operation process is completed, integrate the written fields into the complete structure according to the nodes and output the tree structure configuration entity that can be called, generating the backup index tree configuration result.

[0033] The cloud backup system of the communication software is used to execute the above-mentioned cloud backup method of the communication software. The system includes: The node access and collection module extracts the device number, node address, and access time fields based on the cloud backup nodes connected by users in the communication software, groups them according to the device number, arranges the node order according to the access time, judges the continuous access node jump, and generates a device behavior path mapping table; The behavior path mapping module uses the device behavior path mapping table to identify the node combinations in the jump sequence, counts the occurrence frequency of the node combinations, and combines with the device number to obtain the cloud backup scheduling node chain group; The node label screening module calls the cloud backup scheduling node chain group, extracts the client labels bound to the associated data items in sequence based on the path order, and generates a unified session data time table by comparing the node labels with the data item labels one by one; The session time unification module calls the communication records of the channels in the unified session data time table, extracts the device number, session identifier, and message sending time, classifies the data according to the device number, locates the sending time of each session, and generates a field order index path set; The backup index module calls the field order index path set, extracts the sending time, message length, file format, and session number, sorts the fields based on the sending time, extracts the field path positions, constructs index nodes in the path order, binds the data records, and obtains the backup index tree configuration result.

[0034] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A cloud backup method for a communication software, characterized in that, It includes the following steps: S1: Based on the cloud backup nodes connected by users in the communication software, collect the access log data of the nodes on different devices, extract the node jump sequences in the continuous access behaviors according to the device numbers, identify the repeated node combinations, and obtain the device behavior path mapping table; S2: Call the device paths in the device behavior path mapping table, sequentially detect the client tags bound in the associated data items according to the path order, compare the tags of the data items with the nodes in the path, screen the node data combinations with the behavior types consistent with the path order, and output the cloud backup scheduling node chain group; S3: Call the communication records in each channel in the cloud backup scheduling node chain group, extract the device number, session identifier and message sending time, classify by device and then locate the start node time field of the session, perform time comparison and then uniformly revise the time field of the device record to obtain the unified session data time table; S4: Call the field information of each data in the unified session data time table, extract the sending time, message length, file format and session number fields, combine with the access order of the corresponding fields, and arrange the field structure according to the access sequence to form the field order index path set.

2. The cloud backup method of the communication software according to claim 1, wherein The device behavior path mapping table includes a path number, a node jump mode, and a device behavior tag. The cloud backup scheduling node chain group includes a channel identifier, a node binding mapping, and a path sequence number. The unified session data time table includes a message time mapping table, a device unified timestamp, and a synchronization status identifier. The field order index path set includes a field order list, a field position index, and an access structure tag.

3. The cloud backup method of the communication software according to claim 1, characterized in that, The specific steps for obtaining the device behavior path mapping table are as follows: S111: Based on the cloud backup nodes connected by users in the communication software, collect the access log data on different devices, extract the device number and the complete sequence of the corresponding access behaviors, and arrange the log records in the order of access to obtain the device access behavior sequence group; S112: Call the device access behavior sequence group, for the continuous access behavior sequences under the same device number, extract the node number fields in the access records, compare the node jump paragraphs, identify the jump segments with the same node combinations, and perform number mapping on the repeated jump segments to obtain the node jump combination set; S113: According to the node jump combination set, accumulate the occurrence frequencies of the repeated combinations in the node jump paths of each device, combine the node jump path length and the position of the jump combination in the path, calculate the mapping intensity value of the device in the node combination path, summarize the mapping intensity values by device number respectively, analyze the node behavior relationship of the device, and obtain the device behavior path mapping table.

4. The cloud backup method of the communication software according to claim 3, wherein The specific steps for obtaining the cloud backup scheduling node chain group are as follows: S211: Call the device paths in the device behavior path mapping table, sequentially detect the client tags bound in the associated data items according to the path order, compare the tags of the data items with the nodes in the path, and screen the node data combinations according to the consistency of the data item tags and the node tags to obtain the node data combinations with consistent tags; S212: Based on the tag-consistent node data combination, filter the node data combinations with consistent behavior types and path orders, calculate the behavior order matching degree value, filter the node data combinations that meet the requirements according to the behavior order matching degree value, and generate the order-matching node data combinations; S213: Based on the order-matching node data combinations, combine the continuous mapping relationship between nodes and data items, classify the node chain groups that meet the mapping relationship into the scheduling channels, and output the cloud backup scheduling node chain groups.

5. The cloud backup method of the communication software according to claim 4, characterized in that, The specific steps for obtaining the unified session data time table are as follows: S311: Based on the communication records of each channel in the cloud backup scheduling node chain groups, extract the device number, session identifier, and message sending time, classify them according to different device numbers, and aggregate the communication records with the same device number to obtain the device communication record set; S312: Invoke the device communication record set, for each session identifier in the set, locate the start node time field of the session, extract the start node time field data, compare the start node times under different device numbers, calculate the unified time difference metric value between devices, uniformly revise the device record time fields, and integrate the time records of each session identifier in combination with the revised time field set to generate the unified session data time table.

6. The cloud backup method of the communication software according to claim 5, wherein The specific steps for obtaining the field order index path set are as follows: S411: Invoke the field information of multiple pieces of data in the unified session data time table, extract the fields of the sending time, message length, file format, and session number recorded in each piece of data, and for each piece of data, extract the corresponding fields in sequence according to the field identification position to obtain the field combination set; S412: Use the content of the fields of the sending time, message length, file format, and session number in the field combination set, combine the field access order information recorded in the client access log data, compare it with the field access order number under the same number in the access log, and combine the real-time arrangement sequence of each group of data fields to obtain the field order index path set.

7. The cloud backup method of the communication software according to claim 1, characterized in that, The method further includes the step: S5: Invoke the path positions corresponding to the fields in the field order index path set, construct multi-level index nodes in sequence, complete the configuration of the upper, middle, and lower layer nodes according to the path order, bind each data write request to the corresponding structure path, and generate the backup index tree configuration result; The backup index tree configuration result includes the index node structure, data binding relationship, and structure level identifier.

8. The cloud backup method of the communication software according to claim 7, wherein The specific steps for obtaining the backup index tree configuration result are as follows: S511: Invoke the path positions corresponding to the fields in the field order index path set, and according to the field arrangement order, locate the position information of the fields in the path set in sequence, and combine the position information with the field names to generate the field path index order; S512: Based on the field path index order, construct multi-level index nodes in sequence, aggregate the fields with the same path positions in the same level into the same node, and respectively set the structural relationships of the upper, middle, and lower layers to generate the index level node connection structure; S513: Invoke the index level node connection structure to perform path recognition processing on each data write request, bind the data content to the corresponding node positions according to the field mapping relationship, and sequentially nest and construct a tree structure according to the structure level to generate the backup index tree configuration result.

9. A cloud backup system for a communication software, characterized in that, The system is used to implement the cloud backup method of the communication software according to any one of claims 1-8, and the system includes: The node access and collection module extracts the device number, node address, and access time fields based on the cloud backup nodes connected by the user in the communication software, groups them according to the device number, arranges the node order according to the access time, determines the continuous access node jump, and generates a device behavior path mapping table; The behavior path mapping module uses the device behavior path mapping table to identify the node combinations in the jump sequence, counts the occurrence frequency of the node combinations, and combines with the device number to obtain the cloud backup scheduling node chain group; The node label screening module invokes the cloud backup scheduling node chain group, sequentially extracts the client labels bound to the associated data items based on the path order, and generates a unified session data time table by comparing the node labels with the data item labels one by one; The session time unification module invokes the communication records of the channels in the unified session data time table, extracts the device number, session identifier, and message sending time, classifies the data according to the device number, locates the sending time of each session, and generates a field order index path set; The backup index module invokes the field order index path set, extracts the sending time, message length, file format, and session number, sorts the fields based on the sending time, extracts the field path positions, constructs index nodes according to the path order, and binds the data records to obtain the backup index tree configuration result.

Citation Information

Patent Citations

  • Method and system for data increment backup of sensing layer of Internet of Things

    CN102981933A

  • Remote data backup method and device and computer readable medium

    CN108459926A

  • Cloud storage data synchronization method and device and storage medium

    CN119046377A

  • Network protocol information retrieval method and device, computer equipment, readable storage medium and program product

    CN119226327A

  • Masterless backup and restore of files with multiple hard links

    US20200250141A1