Memory resource usage control method and system for Java data analysis software
By logically grouping and managing memory resources of Java data analysis software, designing columnar storage structures and memory usage statistics, and combining neural network prediction models, the problems of inaccurate memory usage and difficult resource isolation in Java data analysis software are solved, and efficient memory resource management and resource isolation of analysis tasks are achieved.
Patent Information
- Application Number
- CN202510134196.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The prior art is difficult to achieve accurate statistics and regulation control of memory resources in Java data analysis software, resulting in inaccurate memory usage and inability to achieve resource isolation.
By logically grouping the memory resources of Java data analysis software and managing them based on the group memory manager, columnar storage structures of different data types are designed, memory usage of operators is calculated, comprehensive memory usage sequences are generated, and memory overflow detection is carried out through neural network prediction models to achieve resource isolation.
It improves the accuracy and control efficiency of memory resource use, realizes resource isolation between different analysis tasks, avoids memory competition and overflow, and improves the operation efficiency of data analysis software.
Smart Images

Figure CN119576588B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of resource allocation control, and in particular to a method and system for controlling memory resource usage of Java data analysis software. Background Art
[0002] Software developed using the Java language generally runs on a Java virtual machine (JVM). The memory resource management mechanism of the JVM is relatively simple: first, after the JVM is started, it applies for memory resources from the operating system. Then, when the application running on the JVM creates an object instance, memory is allocated according to the object definition. Finally, the JVM's built-in garbage collector scans the entire object instance created by the JVM, and starts to reclaim this part of the memory space when it finds no referenced objects. This design will allow users who use the JVM to not worry about the application and release of memory, and likewise lack methods to intervene and control it. At present, some applications bypass the JVM by directly applying for system memory, but this is limited to local or simple scenarios and cannot be fully implemented. Therefore, it is necessary to provide a new memory usage control method and system to solve the above problems.
[0003] Similar prior art includes a Chinese patent application with publication number CN118394602A, which discloses a memory calculation method, device and electronic device for a Java object, and relates to the field of computer technology. The method determines the corresponding target memory calculation method according to the data type of the Java object to be calculated, and then uses the target memory calculation method to obtain the memory occupied by the Java object to be calculated. The scheme can select different memory calculation methods according to the data type of the Java object to realize memory calculation, and can more accurately calculate the memory occupied by the Java object, and then can realize a more fine-grained detection of the memory occupied by the Java object inside the program, which is conducive to the program to accurately understand its own memory occupation. There is also a Chinese patent application with publication number CN112835765A, which discloses a Java program memory usage monitoring method and device, which can be applied to the financial field. The method includes: obtaining the indicator data of each area in the Java memory in real time by calling the Java management extension interface, and the indicator data includes: memory usage, maximum memory usage, garbage collection times and garbage collection time; aggregating the indicator data according to the set time unit through the real-time calculation database to obtain the memory usage data of each area; and providing real-time warning for the memory usage of each area based on the size relationship between the memory usage data of each area and the preset threshold. The Java program memory usage monitoring method provided by the application partitions the memory, and then monitors the Java program memory for each partition in combination with the garbage collection situation and memory occupancy, achieving the technical effect of accurately monitoring the memory without affecting the system performance.
[0004] The shortcomings of the existing technology are mainly manifested in that different memory calculation methods are selected only through the data type of the Java object to realize memory calculation, which will lead to inaccurate statistics of system memory resources. It only monitors the memory of the Java program and does not adjust and control the memory resources. Summary of the invention
[0005] The present application provides a method and system for controlling memory resource usage of Java data analysis software, which are used to improve the efficiency and accuracy of memory resource usage control of Java data analysis software.
[0006] In a first aspect, the present application provides a method for controlling memory resource usage of a Java data analysis software, the method comprising:
[0007] Logically grouping the memory resources of the analysis software to generate multiple groups, managing the memory occupancy of the group based on the group memory manager, setting the memory corresponding to the group, the analysis software obtaining the group information of the group after receiving the request, transmitting the group information based on the execution environment and thread context, and maintaining the memory of the group based on the group memory manager;
[0008] Based on different data types, a column storage structure corresponding to the group Group is defined, the column storage structure obtains a current memory size based on a defined variable, and a first memory occupancy of the column storage structure is counted based on the memory size;
[0009] Counting the second memory usage of any operator during the calculation process, and combining the second memory usage based on the number of types of the operators to generate the calculation memory used by the analysis software;
[0010] The first memory occupancy is combined with the computing memory to form a comprehensive memory occupancy, a sampling frequency is set, and the comprehensive memory occupancy of the analysis software when it is running is collected based on the sampling frequency to generate a memory occupancy sequence, and the group memory manager performs memory overflow detection based on the memory occupancy sequence;
[0011] If the computing task of the analysis software includes associated computing, and the data volume of the associated computing is greater than a first threshold, the associated computing is set as a computing operation, computing nodes are set, and computing operations are allocated based on the computing nodes to achieve resource isolation.
[0012] In combination with the first aspect, the storage structure obtains the current memory size based on the defined variable, including:
[0013] The column storage structure has a unified add(i) interface and get(i) interface, wherein i represents a position index number, the defined variable is initialized to 0, and the defined variable is updated after adding the data to be analyzed in the add(i) interface;
[0014] The data types include general numerical types, high-precision numerical types, string types, collection types, and object types. If the data to be analyzed is of the general numerical type, it is stored in the column storage structure based on the primary type of Java, and the product of the number of elements of the data to be analyzed and the memory occupied by a single element is set as the current memory size;
[0015] If the data to be analyzed is of the high-precision numerical type, a fixed-precision data type FixedDecimal is defined in the column storage structure based on the applicable characteristics of the high-precision numerical type as a replacement, and the memory occupancy of the internal fixed constant level is set to the current memory size;
[0016] If the data to be analyzed is of the string type, the String type of JDK is used in the column storage structure, a string pool is set, the string value corresponding to the data to be analyzed is input into the string pool, and the position information of the string value is stored. The string pool is shared within the group Group, and the memory occupied by the string pool is set to the current memory size;
[0017] If the data to be analyzed is of the set type, different types of set types are established in the column storage structure, a set interface is set, and based on the set interface, memory occupancy of different types of set types is obtained and set as the current memory size;
[0018] If the data to be analyzed is of the object type, the object supports serialization and deserialization to an interface of a byte array in the column storage structure, the serialized byte data of the object is stored, and the memory occupancy of the byte array after storage is counted based on the binary data and set as the current memory size.
[0019] In combination with the first aspect, based on the type of the data type included in the data to be analyzed, all the memory sizes are summarized and counted to generate the first memory occupancy.
[0020] In combination with the first aspect, the operator includes a filtering operator, a conversion operator, an association operator, a grouping operator, a summary operator, a sorting operator, and a combination operator after any multiple of the operators are combined, and the second memory usage of any operator in the calculation process is counted, including:
[0021] If the operator is the filter operator, the storage added during the calculation process is the first integer array, and the memory occupied by the first integer array is set to the second memory occupancy;
[0022] If the operator is the conversion operator, the storage added during the calculation process is a new column memory, and the memory occupied by the new column memory is counted and set as the second memory occupancy;
[0023] If the operator is the association operator, the storage added during the calculation process is the Data Holder generated by scanning the association column of the right table and the second integer array identifying the position of the association key of the right table in the Data Holder, and the memory occupied by the Data Holder and the second integer array is set as the second memory occupation;
[0024] If the operator is the grouping operator, the storage added during the calculation process is the storage of the grouping key value, and the memory occupied by the storage of the grouping key value is set to the second memory occupancy;
[0025] If the operator is the summary operator, the storage added during the calculation process is the intermediate result of the summary, and the memory occupied by the intermediate result is set to the second memory occupancy;
[0026] If the operator is the sorting operator, the storage added during the calculation process is two integer arrays, and the memory occupied by the two integer arrays is set as the second memory occupancy;
[0027] If the operator is a combination operator, the occupied memory is accumulated based on the operator types included in the combination operator and set as the second memory occupation.
[0028] In combination with the first aspect, combining the second memory occupancy based on the number of types of the operators to generate the computing memory used by the analysis software includes:
[0029] If the number of operator types is 1, the second memory occupancy corresponding to the operator is set as the memory used for operation based on the type of the operator; otherwise, the second memory occupancy corresponding to all the operators is accumulated based on the number of operator types to generate the memory used for operation.
[0030] In combination with the first aspect, the group memory manager performs memory overflow detection based on the memory occupancy sequence, including:
[0031] Summarize the comprehensive memory occupancy corresponding to all objects based on the acquisition time point to generate a total memory occupancy sequence, calculate the ratio of the memory occupancy sequence to the total memory occupancy sequence based on the same acquisition time point, and summarize all the ratios based on the acquisition time point to generate a ratio sequence;
[0032] Building a prediction model based on a neural network, wherein the prediction model predicts the memory occupancy sequence based on the proportion sequence and then outputs a first prediction result;
[0033] The group memory manager checks any of the groups Group, and if the first prediction result is greater than a second threshold, determines that it is abnormal, and terminates the execution of the operator;
[0034] A group container is established, and the group container obtains all objects to be counted based on a weak reference mechanism. When all the objects to be counted in the group container are reclaimed by the garbage collector of the JVM, the group Group is deleted based on the group memory manager, and the memory of the group Group is released and allocated to the remaining group Groups for use.
[0035] In combination with the first aspect, the setting of the computing node further includes:
[0036] The type of the computing node is set based on the hardware device of the analysis software, and a network interface, a storage device and a computing device are configured for the computing node. The supervisory process of the analysis software distributes the application code to all computing nodes participating in the computing operation, obtains the supervisory process of the computing node and generates process instructions. The process of the computing node generates computing operations based on the process instructions, and the computing device executes the computing operations.
[0037] In combination with the first aspect, allocating computing operations based on the computing nodes includes:
[0038] Based on the storage device, a computing node with the smallest computing load is extracted from all the computing nodes and set as a candidate computing node. All parameter information required for the computing operation is sent to the candidate computing node based on the RPC method. The candidate computing node completes the calculation and outputs the calculation result. The calculation result is sent to the current computing node based on the RPC method. This step is repeated until the computing operation is completed.
[0039] In combination with the first aspect, the transmitting the grouping information based on the execution environment and the thread context includes:
[0040] Creating a corresponding execution environment for each analysis task, the execution environment receiving a processing instruction corresponding to the grouping information, and the analysis software acquiring a data set and an analysis type of the grouping information based on the processing instruction;
[0041] Based on the thread manager, multiple threads are set, the data set and the analysis type are passed to the thread context based on the thread parameters, the multiple threads are executed in parallel, each thread processes one data set, and the thread results are output based on the analysis type, and the thread results of all the threads are aggregated to generate the analysis result.
[0042] In a second aspect, the present application provides a memory resource usage control system for Java data analysis software, the memory resource usage control system for Java data analysis software comprising:
[0043] A logical grouping module is used to logically group the memory resources of the analysis software, generate multiple groups, manage the memory usage of the group based on the group memory manager, set the memory corresponding to the group, the analysis software obtains the group information of the group after receiving the request, transmits the group information based on the execution environment and thread context, and maintains the memory of the group based on the group memory manager;
[0044] A first statistics module is used to define a column storage structure corresponding to the group Group for different data types, the column storage structure obtains a current memory size based on a defined variable, and counts a first memory occupancy of the column storage structure based on the memory size;
[0045] A second statistical module is used to count the second memory usage of any operator during the calculation process, and to combine the second memory usage based on the number of types of the operators to generate the calculation memory used by the analysis software;
[0046] a detection module, configured to combine the first memory occupancy with the computing memory to set the combined memory occupancy, set a sampling frequency, collect the combined memory occupancy of the analysis software when it is running based on the sampling frequency to generate a memory occupancy sequence, and the group memory manager performs memory overflow detection based on the memory occupancy sequence;
[0047] The resource isolation module sets the associated calculation as a calculation operation, sets the calculation node, and allocates the calculation operation based on the calculation node to achieve resource isolation if the calculation task of the analysis software includes associated calculation and the data volume of the associated calculation is greater than a first threshold.
[0048] In the technical solution provided by the present application, by logically grouping memory resources and managing them based on a group memory manager, resource isolation between different analysis tasks is achieved, and memory competition and mutual interference between different tasks are avoided. Different column storage structures are designed for different data types, the current memory size is obtained based on defined variables, and different memory occupancy calculation methods are used for different data types. The current memory size is obtained based on defined variables, and different memory occupancy calculation methods are used for different data types, which effectively reduces the memory occupancy and realizes the accurate calculation of memory occupancy. Not only the first memory occupancy of the column storage structure is counted, but also the second memory occupancy of the calculator during the calculation process is counted, and they are combined to generate a comprehensive memory occupancy, realizing a comprehensive statistics of memory occupancy.
[0049] This application collects and analyzes the comprehensive memory occupancy of any object during the operation of the software based on the sampling frequency, generates a memory occupancy sequence, realizes real-time monitoring of memory usage, and through the analysis of the memory occupancy sequence, can understand the changing trend of memory usage, and provide data support for memory overflow prediction. A neural network is also used to build a prediction model, predict the memory occupancy based on the memory occupancy sequence, and output the first prediction result, which realizes the accurate prediction of the memory overflow risk, can prevent the memory occupancy from further increasing, and avoids the occurrence of memory overflow. Finally, this application realizes the optimal configuration and efficient utilization of computing resources and improves the operating efficiency of data analysis by means of computing node type setting and hardware configuration, load balancing and computing operation allocation, multi-threaded parallel processing and result aggregation. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.
[0051] Figure 1 A schematic diagram of an embodiment of a method for controlling memory resource usage of Java data analysis software in an embodiment of the present application;
[0052] Figure 2 This is a schematic diagram of an embodiment of a memory resource usage control system of Java data analysis software in an embodiment of the present application. DETAILED DESCRIPTION
[0053] The embodiment of the present application provides a method and system for controlling the use of memory resources of a Java data analysis software. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0054] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 , an embodiment of the memory resource usage control method of Java data analysis software in the embodiment of the present application includes:
[0055] S101. Logically group the memory resources of the analysis software to generate multiple groups. Manage the memory usage of the groups based on the group memory manager. Set the memory corresponding to the groups. After the analysis software receives the request, it obtains the group information of the groups. The group information is transmitted based on the execution environment and thread context. The memory of the groups is maintained based on the group memory manager.
[0056] It is understandable that the execution subject of the present application can be a memory resource usage control system of the Java data analysis software, or a terminal or a server, which is not limited here. The present application embodiment is described by taking the server as the execution subject as an example.
[0057] Specifically, the analysis software refers to the system corresponding to the Java data analysis software. Logical grouping refers to logical division, not physical division. This grouping can be a query calculation, a user's access, or a tenant's use. For example, according to the user's requirements for memory usage (such as high-priority tasks, low-priority tasks, etc.), real-time analysis tasks that require fast response can be assigned to the "high priority" group, and batch tasks can be assigned to the "low priority" group. Grouping is similar to the concept of control groups in the Linux operating system. Grouping represents the basic unit for resource application. Group information is determined after the analysis software application receives the request (because the request contains the calculation to be performed, user information, and tenant information), and the group can be registered through the group memory manager. After registration, the memory usage of the group can be managed in the group memory manager, including the add, remove, and used methods to increase the memory usage, reduce the memory usage, and query the memory usage. The maximum available memory of a group can be configured through the application configuration of the analysis software, but the upper limit cannot exceed half of the maximum allocated memory of the JVM to avoid affecting normal business processing. Group information can be transmitted and shared through the execution environment and thread context to ensure that the group to which it belongs can be obtained at any time during the operation, and the memory usage is maintained through the group memory manager. For example, maintenance functions such as memory allocation and recycling, memory fragmentation management, and memory leak detection.
[0058] S102: Define a column storage structure corresponding to the group Group based on different data types, obtain a current memory size based on a defined variable, and count a first memory occupancy of the column storage structure based on the memory size.
[0059] Specifically, most data analysis applications in analysis software use column storage, and corresponding column storage structures are defined for different data types, and different storage structures have the characteristic of convenient memory occupancy statistics. The acquisition of memory size includes but is not limited to static analysis and dynamic calculation, wherein static analysis refers to estimating the memory occupancy of each column storage structure through static code analysis tools during the compilation phase, and dynamic calculation refers to dynamically calculating the memory occupancy of each column storage structure according to the actual data volume during runtime. The first memory occupancy refers to the memory space occupied by the column storage structure itself, excluding the memory temporarily occupied during data processing. Since the group Group may contain data of multiple data types to be analyzed, it is necessary to count the memory size to generate the first memory occupancy. By defining a column storage structure, data can be compressed to reduce memory occupancy, and efficient encoding methods, such as dictionary encoding, bitmap encoding, etc., can be used to improve data storage and reading efficiency, and memory alignment can be optimized to reduce memory access conflicts and improve cache hit rate.
[0060] S103: Count the second memory occupancy of any operator during the calculation process, and combine the second memory occupancy based on the number of operator types to generate the calculation memory used by the analysis software.
[0061] Specifically, in the field of data analysis, operators usually refer to functions or methods that perform certain operations on data. For example, when performing data analysis in Java, operators can be functions that clean, transform, aggregate, and perform other operations on data. In this application, operators include but are not limited to filtering, conversion, association, grouping, aggregation, sorting, or a combination of such operators. The second memory usage refers to the memory space temporarily occupied by the operator during the calculation process, including intermediate results, temporary variables, etc., which is a memory usage statistics of the processing objects in the operator calculation process. The number of types refers to the total number of operators of various types, and the memory used for operation is the combination of the second memory usage of all operators to obtain the total memory used for operation of the analysis software during the data processing process.
[0062] S104, combining the first memory occupancy and the computing memory usage to set as a comprehensive memory occupancy, setting a sampling frequency, collecting and analyzing the comprehensive memory occupancy of the software at runtime based on the sampling frequency to generate a memory occupancy sequence, and the group memory manager performs memory overflow detection based on the memory occupancy sequence.
[0063] Specifically, the comprehensive memory usage refers to the first memory usage corresponding to the column storage structure in any group Group, and the memory usage after the operation memory usage corresponding to all operators is combined, which can be used to obtain the comprehensive memory usage of the analysis software during operation. The sampling frequency refers to sampling once every certain data unit (the default is 1000, which can be flexibly modified through configuration), where the number of data units = the total number of columns Scan the number of rows. When calculating, the operator scans the data source (using column storage), scans the data row by row, and counts the memory usage of the runtime object to avoid performance overhead. The comprehensive memory usage at any collection time point is arranged in sequence according to the collection time point corresponding to the collection frequency to generate a memory usage sequence. Memory overflow detection refers to determining whether there is a risk of memory overflow. For example, when the memory usage continues to grow and approaches the memory limit, it is determined that the risk of memory overflow is high. In addition, memory processing methods can be made based on memory overflow detection.
[0064] S105. If the computing task of the analysis software includes associated computing, and the data volume of the associated computing is greater than a first threshold, the associated computing is set as a computing operation, computing nodes are set, and computing operations are allocated based on the computing nodes to achieve resource isolation.
[0065] Specifically, associated computing refers to computing tasks that have associated operators in the operator type. Such computing tasks usually require special processing methods due to their large amount of computing and high memory consumption, such as isolating them to run in independent JVM processes to ensure the stability of the overall system performance and efficient use of resources. Computing operations refer to computing tasks corresponding to associated computing. Resource isolation refers to allocating independent memory resources to computing operations to achieve resource isolation and avoid competing with other tasks for memory resources.
[0066] In a specific embodiment, the storage structure obtains the current memory size based on the defined variables, including:
[0067] (1) The column storage structure has a unified add(i) interface and get(i) interface, where i represents the position index number. The defined variable is initialized to 0. After the data to be analyzed is added in the add(i) interface, the defined variable is updated.
[0068] (2) Data types include general numeric types, high-precision numeric types, string types, collection types, and object types. If the data to be analyzed is of a general numeric type, it is stored in a columnar storage structure based on Java's primary type, and the product of the number of elements in the data to be analyzed and the memory occupied by a single element is set to the current memory size.
[0069] (3) If the data to be analyzed is a high-precision numeric type, a fixed-precision data type FixedDecimal is defined in the column storage structure based on the applicable characteristics of the high-precision numeric type as a replacement, and the memory usage of the internal fixed constant level is set to the current memory size.
[0070] (4) If the data to be analyzed is of string type, the JDK String type is used in the column storage structure, a string pool is set, the string value corresponding to the data to be analyzed is input into the string pool, and the location information of the string value is stored. The string pool is shared within the group Group, and the memory occupied by the string pool is set to the current memory size.
[0071] (5) If the data to be analyzed is of a collection type, different types of collection types are established in the column storage structure, a collection interface is set, and the memory usage of different types of collection types is obtained based on the collection interface and set as the current memory size.
[0072] (6) If the data to be analyzed is of object type, the object is supported to be serialized and deserialized to a byte array interface in the column storage structure, the serialized byte data of the object is stored, and the memory usage of the byte array after storage is counted based on the binary data and set as the current memory size.
[0073] Specifically, the present application further elaborates on the method of obtaining the memory size of different data types in the column storage structure. The add(i) interface is used to add the data to be analyzed to the column storage structure. i represents the position index where the data is added. After the data is added, the defined variable used to track the memory size is updated. The get(i) interface is used to obtain the data at a specified position from the column storage structure. i represents the position index where the data is obtained. A variable is defined in the column storage structure to save the memory size occupied by the current structure. The variable is updated when add is completed. The defined variable is initialized to 0 and is used to record the current memory size through update accumulation.
[0074] Generally, numeric types refer to date types, Boolean types, and data types such as int, long, float, and double in numeric types. The calculation is based on the number of bytes of primitive types (i.e., primary types) in Java, for example, int is 4 bytes, long is 8 bytes, float is 4 bytes, and double is 8 bytes.
[0075] High-precision numeric types refer to BigDecimal, BigInteger and other data types. Since their memory usage is large and not fixed, direct storage will lead to memory waste. In addition, the data type FixedDecimal has a fixed memory usage. Therefore, the data type FixedDecimal is used to replace high-precision numeric types. The memory usage of the internal fixed constant level of the data type FixedDecimal is set to the current memory size. For example, assuming that each instance of the data type FixedDecimal occupies 128 bytes, the current memory size is equal to 128.
[0076] In the case of string types, in order to save memory, the string pool mechanism is used to store all string values in the string pool. The column storage structure only stores the location information of the string value in the pool. The string pool is shared within the same group Group to avoid repeated storage of the same string value. In the add(i) interface, check whether the same string value already exists in the string pool. If it does, get its location information and update the definition variable (only increase the reference memory usage). If it does not exist, add the string value to the string pool, get its location information, and update the definition variable (increase the string memory usage and reference memory usage). The memory occupied by the string pool refers to the total memory usage of all strings in the string pool, including string content, references, etc.
[0077] The situation of collection type data is more complicated. It mainly supports the following common collection types: List, Map, and Set. The collection interface is used to obtain the memory usage of different types of collection types. For example, the collection interface Collection <t>The current memory size can be generated by adding up the memory usage of different types of collections.
[0078] For object types, in the add(i) interface, the object is serialized into a byte array, and the memory usage of the byte array can be calculated as the memory size.
[0079] In a specific embodiment, based on the types of data included in the data to be analyzed, all memory sizes are aggregated and counted to generate a first memory occupancy.
[0080] Specifically, the data to be analyzed may have multiple data types. Corresponding memory sizes are obtained according to different data types, and then all memory sizes are aggregated to generate the first memory occupancy.
[0081] Through the above implementation, the present application can accurately obtain the memory usage of different data types and incorporate it into the comprehensive memory usage calculation, thereby achieving refined management of the memory resource usage of Java data analysis software.
[0082] In a specific embodiment, the operator includes a filter operator, a conversion operator, a correlation operator, a grouping operator, a summary operator, a sorting operator, and a combination operator formed by combining any multiple operators, and counting the second memory usage of any operator during the calculation process includes:
[0083] (1) If the operator is a filter operator, the storage added during the calculation process is the first integer array, and the memory occupied by the first integer array is set to the second memory occupancy.
[0084] (2) If the operator is a conversion operator, the storage added during the calculation process is a new column memory. The memory occupied by the new column memory is counted and set as the second memory occupancy.
[0085] (3) If the operator is an associative operator, the storage added during the calculation process is the Data Holder generated by scanning the associative column of the right table and the second integer array that identifies the position of the associative key of the right table in the Data Holder. The memory occupied by the Data Holder and the second integer array is set to the second memory occupancy.
[0086] (4) If the operator is a grouping operator, the storage added during the calculation process is the storage of the grouping key value, and the memory occupied by the storage of the grouping key value is set as the second memory occupancy.
[0087] (5) If the operator is a summary operator, the storage added during the calculation process is the intermediate result of the summary, and the memory occupied by the intermediate result is set as the second memory occupancy.
[0088] (6) If the operator is a sorting operator, the storage added during the calculation process is two integer arrays, and the memory occupied by the two integer arrays is set as the second memory occupancy.
[0089] (7) If the operator is a combination operator, the occupied memory is accumulated based on the operator types contained in the combination operator and set as the second memory occupancy.
[0090] Specifically, the present application further elaborates on how to count the second memory usage of different types of operators during the calculation process, and how to generate the calculation memory usage of the analysis software based on the number of operator types.
[0091] The filter operator is used to filter data according to the specified conditions, retaining the records that meet the conditions and removing the records that do not meet the conditions. The first integer array is used to store the row indexes that meet the filter conditions. For example, assuming that the original data has 1000 rows and 500 rows remain after filtering, the size of the first integer array is 500, and the second memory occupancy = the length of the first integer array × the memory occupied by the integer. For example, assuming that the integer occupies 4 bytes, the second memory occupancy = 500 × 4 = 2000 bytes. During the execution of the filter operator, the original data is traversed, and each record is judged according to the filter conditions to see whether it meets the conditions. The row indexes that meet the conditions are stored in the first integer array, the memory occupancy of the first integer array is calculated, and it is set as the second memory occupancy.
[0092] The conversion operator is used to perform some form of conversion on the data, such as type conversion, value mapping, etc. The new column storage is used to store the converted data. For example, if the int type is converted to the String type, the new column storage is used to store the converted string value. The second memory usage = the length of the new column storage × the memory occupied by a single element. For example, assuming that the average length of the converted string is 20 bytes, the second memory usage = data length × 20. During the execution of the conversion operator, the original data is traversed, new data is generated according to the conversion rules, the converted data is stored in the new column storage, the memory usage of the new column storage is calculated, and it is set as the second memory usage.
[0093] The association operator is used to connect two data sets according to the specified key, such as left join, right join, inner join, etc. The Data Holder generated by scanning the association column of the right table is used to store the data obtained by scanning the association column of the right table (the table with fewer rows in the two associated tables). The second integer array that identifies the position of the right table association key in the Data Holder is used to store the position index of the right table association key in the Data Holder. The second memory usage = Data Holder memory usage + second integer array memory usage. For example, assuming that the Data Holder occupies 10,000 bytes and the length of the second integer array is 500, the second memory usage = 10,000 + 500 × 4 = 12,000 bytes. During the execution of the association operator, the right table association column is scanned, the data is stored in the Data Holder, the position index of each right table association key in the Data Holder is recorded and stored in the second integer array, the memory usage of the Data Holder and the second integer array is calculated, and it is set as the second memory usage.
[0094] The grouping operator is used to group data according to the specified key for subsequent aggregation operations. The storage grouping key value is used to store the value of the grouping key. The second memory usage = the storage length of the grouping key value × the memory occupied by a single element. For example, assuming that the grouping key is of type String and the average length is 20 bytes, the second memory usage = the number of grouping key values × 20. During the execution of the grouping operator, the data is traversed, grouped according to the grouping key value, stored in the memory, the memory usage of the grouping key value is calculated, and set as the second memory usage.
[0095] The summary operator is used to perform aggregation operations on the grouped data, such as sum, count, average, etc. The intermediate results of the summary are used to store the intermediate results of the summary operation. The second memory usage = the number of intermediate results × the memory occupied by a single intermediate result. For example, assuming there are 1000 groups and each intermediate result occupies 8 bytes, the second memory usage = 1000 × 8 = 8000 bytes. During the execution of the summary operator, the grouped data is summarized, the intermediate results of the summary are stored in the memory, the memory usage of the intermediate results is calculated, and it is set as the second memory usage.
[0096] The sorting operator is used to sort data. In the two integer arrays, the first integer array is used to store the index of the sorted data, and the second integer array is used to store the temporary index used in the sorting process. The second memory usage = (2 × data length × integer array memory usage). For example, assuming that the data length is 1000 and the integer array occupies 4 bytes, the second memory usage = 2 × 1000 × 4 = 8000 bytes. During the execution of the sorting operator, the sorting algorithm is used to sort the data, and two integer arrays are used to store the index information in the sorting process. The memory usage of the two integer arrays is calculated and set as the second memory usage.
[0097] The combination operator refers to a set of types included in multiple operators. Through the execution process of the above operators, the corresponding occupied memories are accumulated in sequence to generate the second memory occupation.
[0098] The data and data sets mentioned above refer to the data to be analyzed contained in the group Group.
[0099] In a specific embodiment, combining the second memory usage based on the number of operator types to generate the computing memory used by the analysis software includes:
[0100] If the number of operator types is 1, the second memory occupancy corresponding to the operator is set as the computing memory based on the operator type; otherwise, the second memory occupancy corresponding to all operators is accumulated based on the number of operator types to generate the computing memory.
[0101] Specifically, after counting the second memory usage of all operators, it is necessary to combine them according to the number of operator types to generate the calculation memory used by the analysis software. If the number of operator types is 1, the second memory usage of the corresponding operator is directly set as the calculation memory used, where the number of types refers to the number of different types of operators.
[0102] In a specific embodiment, the group memory manager performs memory overflow detection based on the memory occupancy sequence, including:
[0103] (1) Summarize the comprehensive memory usage of all objects based on the collection time point to generate a total memory usage sequence. Calculate the ratio of the memory usage sequence to the total memory usage sequence based on the same collection time point. Summarize all ratios based on the collection time point to generate a ratio sequence.
[0104] (2) A prediction model is constructed based on a neural network. The prediction model predicts the memory usage sequence based on the proportion sequence and outputs a first prediction result.
[0105] (3) The group memory manager checks any group Group. If the first prediction result is greater than the second threshold, it is determined to be abnormal and the operator execution is terminated.
[0106] (4) A group container is established. The group container obtains all objects to be counted based on a weak reference mechanism. When all objects to be counted in the group container are reclaimed by the JVM's garbage collector, the group is deleted based on the group memory manager, and the memory of the group is released and allocated to the remaining groups for use.
[0107] Specifically, all objects refer to instances of various data units or data structures that are processed by operators in any group in the analysis software. Proportion calculation refers to calculating the ratio of the subset in the memory usage sequence of each object to the subset of the current total memory usage sequence at each acquisition time point. The proportion data of each object at all acquisition time points are summarized to generate a proportion sequence.
[0108] A prediction model is constructed using a neural network, for example, an LSTM (Long Short-Term Memory) model. The proportion sequence and the memory usage sequence are both input into the prediction model for training and learning. The proportion sequence is used as an influencing factor to predict the change trend of the memory usage sequence in the future, and the first prediction result, i.e., the predicted total memory usage numerical change sequence, is output.
[0109] The group memory manager checks each group and calculates the first prediction results of all objects in each group. If the first prediction results exceed the second threshold, the group is determined to have a memory overflow risk and is considered abnormal. When a group is determined to be abnormal, the group memory manager immediately terminates the operators being executed in the group to prevent further increase in memory usage.
[0110] In order to manage memory resources more effectively, this application introduces the concept of group container and adopts weak reference mechanism for object management. The weak reference mechanism means that the group container obtains all objects to be counted based on the weak reference mechanism, and the weak reference object will not prevent the garbage collector from reclaiming the objects it references. The objects to be counted refer to the weak reference objects of all objects stored in the group container, which can be normally reclaimed by the garbage collector of the JVM. Memory release means that when all objects in the group container are reclaimed by the garbage collector, the group memory manager deletes the corresponding group Group. After deleting the group Group, the memory occupied by it is released. The released memory is allocated to the remaining group Groups for use. For example, assuming that group Group1 is deleted, its released memory can be allocated to group Group2 and group Group3 for use. The above measures work together to effectively prevent memory overflow and improve the stability and efficiency of the analysis software.
[0111] In a specific embodiment, setting a computing node further includes:
[0112] The type of computing node is set based on the hardware equipment of the analysis software, and the network interface, storage device and computing device are configured for the computing node. The supervisory process of the analysis software distributes the application code to all computing nodes participating in the computing operation, obtains the supervisory process of the computing node and generates process instructions. The process of the computing node generates computing operations based on the process instructions, and the computing device executes the computing operation.
[0113] Specifically, the supervisor process of the analysis software first detects the configuration of the hardware device running the software, including the number of CPUs, memory capacity, network bandwidth, storage device type, etc. According to the hardware configuration, the computing nodes are divided into different types. For example, high-performance computing nodes: equipped with multi-core CPUs, large-capacity memory, high-speed network interfaces and SSD storage devices, suitable for computing-intensive tasks. Ordinary computing nodes: equipped with standard CPUs, memory, network interfaces and HDD storage devices, suitable for general computing tasks. Storage-optimized computing nodes: equipped with large-capacity storage devices, suitable for storage-intensive tasks. Network interface configuration refers to configuring the network interface for each computing node, including IP address, port number, network protocol, etc., to ensure that all computing nodes can communicate through the network. Storage device configuration refers to configuring the corresponding storage device according to the computing node type. For example, high-performance computing nodes can improve data reading and writing speeds, and storage-optimized computing nodes can improve storage capacity. Computing device configuration refers to configuring computing device resources such as CPU and memory. Different numbers of CPU cores and memory resources are allocated according to the computing node type. The supervisory process distributes the application code to all computing nodes participating in the computing operation, for example, using a distributed file system (such as HDFS) for code distribution. After each computing node is started, it obtains the connection information of the supervisory process and establishes a communication connection. The supervisory process generates process instructions (ProcessInstructions) based on the computing task requirements and sends them to the corresponding computing nodes. The process instructions contain information such as the type, parameters, and execution order of the computing operation.
[0114] In a specific embodiment, the allocation of computing operations based on computing nodes includes:
[0115] Based on the storage device, the computing node with the smallest computing load is extracted from all computing nodes and set as the candidate computing node. All parameter information required for the computing operation is sent to the candidate computing node based on the RPC method. The candidate computing node completes the calculation and outputs the calculation result. The calculation result is sent to the current computing node based on the RPC method. This step is repeated until the computing operation is completed.
[0116] Specifically, in order to achieve load balancing and improve computing efficiency, this application designs a computing operation allocation method based on the RPC (remote procedure call) mechanism. Regularly collect the load information of each computing node storage device, such as I / O read and write rate, disk usage, etc., and calculate the storage device load score of each computing node. Consider the storage device load and other factors (such as CPU usage, memory usage, etc.) to evaluate the load of each computing node. From all computing nodes, select the node with the smallest load as the candidate computing node. Use RPC to send all parameter information required for the computing operation to the candidate computing node. After receiving the parameter information, the candidate computing node performs the computing operation. After the calculation is completed, the calculation result is sent back to the current computing node via RPC until the computing operation is completed.
[0117] In a specific embodiment, the grouping information is transmitted based on the execution environment and the thread context, including:
[0118] (1) Create a corresponding execution environment for each analysis task. The execution environment receives processing instructions corresponding to the grouping information. The analysis software obtains the data set and analysis type of the grouping information based on the processing instructions.
[0119] (2) Multiple threads are set up based on the thread manager. The data set and analysis type are passed to the thread context based on the thread parameters. Multiple threads are executed in parallel. Each thread processes a data set and outputs thread results based on the analysis type. The thread results of all threads are aggregated to generate analysis results.
[0120] Specifically, the present application further elaborates on how to pass grouping information based on the execution environment and thread context, and realize multi-threaded parallel processing, which can improve the efficiency of data processing in the analysis software. The execution environment creates an independent execution environment for each analysis task to isolate the resource usage and execution status between different tasks. The execution environment includes the following main components: the group information processor is used to receive and process grouping information, the data set manager is used to manage the loading, storage and access of data sets, the analysis type manager is used to manage the definition and execution of analysis types, and the thread manager is used to manage the creation, scheduling and life cycle of threads. Processing instructions refer to the generation of corresponding processing instructions by the analysis software after receiving the analysis task submitted by the user. The processing instructions contain the following information: Group Information specifies the data grouping method, such as grouping by time, grouping by category, etc., the data set identifier specifies the data set to be analyzed, such as the database table name, data file path, etc., and the analysis type specifies the type of analysis task, such as statistical analysis, machine learning model training, etc. The packet information processor of the execution environment receives the processing instruction, parses the packet information, data set identifier and analysis type, passes the parsed information to the corresponding component for subsequent processing, and extracts the data set and analysis type. Among them, the data set refers to a set of related data for analysis. It is the object processed and analyzed by the analysis software, including the original data and metadata related to the data (such as data structure, data type, etc.). The analysis type refers to the type or mode of analysis operation performed on the data set, which defines how the data analysis software processes and analyzes the data, for example, statistical analysis, data visualization, etc.
[0121] In order to achieve multi-threaded parallel processing, this application passes the data set and analysis type to the thread context and creates multiple threads for processing. The thread manager is responsible for creating and managing multiple threads, and decides how many threads to create according to the size of the data set and the complexity of the analysis task. Parameter passing refers to passing the data set and analysis type as parameters to each thread. For example, the data set can be split into multiple subsets, and each thread processes a subset. Each thread performs an analysis task, receives the data set and analysis type as parameters, performs analysis operations, and outputs thread results. In order to generate the final analysis results, the output results of all threads need to be aggregated. The result aggregator can be used to receive the results of all threads and perform aggregation operations, such as summing, merging, statistical analysis, etc. The aggregated results are finally processed, for example, formatted output, stored in a database, generated reports, etc., to generate analysis results. Through the combined effect of the above measures, the computing power of multi-core CPUs can be fully utilized, the execution efficiency of data analysis software can be improved, and the isolation and management of different analysis tasks can be achieved.
[0122] The above describes the memory resource usage control method of the Java data analysis software in the embodiment of the present application. The following describes the memory resource usage control system of the Java data analysis software in the embodiment of the present application. Figure 2 In the embodiment of the present application, an embodiment of the memory resource usage control system of the Java data analysis software includes:
[0123] The logical grouping module 201 is used to logically group the memory resources of the analysis software, generate multiple groups, manage the memory usage of the groups based on the group memory manager, set the memory corresponding to the groups, and obtain the grouping information of the groups after the analysis software receives the request, transmit the grouping information based on the execution environment and thread context, and maintain the memory of the groups based on the group memory manager;
[0124] A first statistics module 202 is used to define a column storage structure corresponding to a group Group for different data types, the column storage structure obtains a current memory size based on a defined variable, and counts a first memory occupancy of the column storage structure based on the memory size;
[0125] The second statistical module 203 is used to count the second memory usage of any operator during the calculation process, and to combine the second memory usage based on the number of operator types to generate the calculation memory used by the analysis software;
[0126] The detection module 204 is used to combine the first memory occupancy and the computing memory to set the comprehensive memory occupancy, set the sampling frequency, collect and analyze the comprehensive memory occupancy of the software when it is running based on the sampling frequency to generate a memory occupancy sequence, and the group memory manager performs memory overflow detection based on the memory occupancy sequence;
[0127] The resource isolation module 205 sets the associated calculation as a calculation operation, sets the calculation node, and allocates the calculation operation based on the calculation node to achieve resource isolation if the calculation task of the analysis software includes associated calculation and the data volume of the associated calculation is greater than the first threshold.
[0128] Through the collaboration of the above components, memory resources are logically grouped and managed based on the group memory manager, which realizes resource isolation between different analysis tasks and avoids memory competition and mutual interference between different tasks. Different column storage structures are designed for different data types, and the current memory size is obtained based on the defined variables. Different memory usage calculation methods are used for different data types. The current memory size is obtained based on the defined variables, and different memory usage calculation methods are used for different data types, which effectively reduces the memory usage and realizes the accurate calculation of memory usage. Not only the first memory usage of the column storage structure is counted, but also the second memory usage of the operator during the calculation process is counted, and they are combined to generate the comprehensive memory usage, realizing the comprehensive statistics of memory usage.
[0129] This application collects and analyzes the comprehensive memory occupancy of any object during the software runtime based on the sampling frequency, generates a memory occupancy sequence, realizes real-time monitoring of memory usage, and through the analysis of the memory occupancy sequence, can understand the changing trend of memory usage, and provide data support for memory overflow prediction. A neural network is also used to build a prediction model, predict the memory occupancy based on the memory occupancy sequence, and output the first prediction result, which realizes the accurate prediction of memory overflow risk, can prevent the memory occupancy from further increasing, and avoids the occurrence of memory overflow. Finally, this application realizes the optimal configuration and efficient utilization of computing resources through means such as computing node type setting and hardware configuration, load balancing and computing operation allocation, multi-threaded parallel processing and result aggregation.
[0130] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0131] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.
[0132] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.< / t>
Claims
1. A method for controlling memory resource usage of Java data analysis software, characterized in that: The memory resource usage control method of the Java data analysis software includes: Logically group the memory resources of the analysis software to generate multiple groups, manage the memory occupancy of the group based on the group memory manager, set the memory corresponding to the group, the analysis software obtains the group information of the group after receiving the request, transmits the group information based on the execution environment and thread context, and maintains the memory of the group based on the group memory manager, wherein the group represents the basic unit of resource application; A column storage structure corresponding to the group Group is defined based on different data types, the column storage structure obtains a current memory size based on a defined variable, and a first memory occupancy of the column storage structure is counted based on the memory size, wherein the first memory occupancy refers to the memory space occupied by the column storage structure itself, excluding memory temporarily occupied during data processing; Counting the second memory occupancy of any operator during the calculation process, and combining the second memory occupancy based on the number of types of the operators to generate the computing memory used by the analysis software, wherein the second memory occupancy refers to the memory space temporarily occupied by the operator during the calculation process; The first memory occupancy is combined with the computing memory to form a comprehensive memory occupancy, a sampling frequency is set, and the comprehensive memory occupancy of the analysis software when it is running is collected based on the sampling frequency to generate a memory occupancy sequence, and the group memory manager performs memory overflow detection based on the memory occupancy sequence; If the computing task of the analysis software includes associated computing, and the data volume of the associated computing is greater than a first threshold, the associated computing is set as a computing operation, computing nodes are set, and computing operations are allocated based on the computing nodes to achieve resource isolation.
2. The memory resource usage control method of Java data analysis software according to claim 1, characterized in that: The storage structure obtains the current memory size based on the defined variables, including: The column storage structure has a unified add(i) interface and get(i) interface, wherein i represents a position index number, the defined variable is initialized to 0, and the defined variable is updated after adding the data to be analyzed in the add(i) interface; The data types include general numerical types, high-precision numerical types, string types, collection types, and object types. If the data to be analyzed is of the general numerical type, it is stored in the column storage structure based on the primary type of Java, and the product of the number of elements of the data to be analyzed and the memory occupied by a single element is set as the current memory size; If the data to be analyzed is of the high-precision numerical type, a fixed-precision data type FixedDecimal is defined in the column storage structure based on the applicable characteristics of the high-precision numerical type as a replacement, and the memory occupancy of the internal fixed constant level is set to the current memory size; If the data to be analyzed is of the string type, the String type of JDK is used in the column storage structure, a string pool is set, the string value corresponding to the data to be analyzed is input into the string pool, and the position information of the string value is stored. The string pool is shared within the group Group, and the memory occupied by the string pool is set to the current memory size; If the data to be analyzed is of the set type, different types of set types are established in the column storage structure, a set interface is set, and based on the set interface, memory occupancy of different types of set types is obtained and set as the current memory size; If the data to be analyzed is of the object type, the object supports serialization and deserialization to an interface of a byte array in the column storage structure, the serialized byte data of the object is stored, and the memory occupancy of the byte array after storage is counted based on the binary data and set as the current memory size.
3. The memory resource usage control method of Java data analysis software according to claim 2, characterized in that: Based on the type of the data type included in the data to be analyzed, all the memory sizes are summarized and counted to generate the first memory occupancy.
4. The memory resource usage control method of Java data analysis software according to claim 1, characterized in that: The operators include filtering operators, conversion operators, association operators, grouping operators, summary operators, sorting operators, and combination operators formed by combining any plurality of the operators, and the statistics of the second memory usage of any operator during the calculation process include: If the operator is the filter operator, the storage added during the calculation process is the first integer array, and the memory occupied by the first integer array is set to the second memory occupancy; If the operator is the conversion operator, the storage added during the calculation process is a new column memory, and the memory occupied by the new column memory is counted and set as the second memory occupancy; If the operator is the association operator, the storage added during the calculation process is the DataHolder generated by scanning the association column of the right table, and the second integer array identifying the position of the association key of the right table in the Data Holder, and the memory occupied by the Data Holder and the second integer array is set to the second memory occupation; If the operator is the grouping operator, the storage added during the calculation process is the storage of the grouping key value, and the memory occupied by the storage of the grouping key value is set to the second memory occupancy; If the operator is the summary operator, the storage added during the calculation process is the intermediate result of the summary, and the memory occupied by the intermediate result is set to the second memory occupancy; If the operator is the sorting operator, the storage added during the calculation process is two integer arrays, and the memory occupied by the two integer arrays is set as the second memory occupancy; If the operator is a combination operator, the occupied memory is accumulated based on the operator types included in the combination operator and set as the second memory occupation.
5. The memory resource usage control method of Java data analysis software according to claim 4, characterized in that: Combining the second memory usage based on the number of types of the operators to generate the computing memory used by the analysis software includes: If the number of operator types is 1, the second memory occupancy corresponding to the operator is set as the memory used for operation based on the type of the operator; otherwise, the second memory occupancy corresponding to all the operators is accumulated based on the number of operator types to generate the memory used for operation.
6. The memory resource usage control method of Java data analysis software according to claim 1, characterized in that: The group memory manager performs memory overflow detection based on the memory occupation sequence, including: Summarize the comprehensive memory occupancy corresponding to all objects based on the acquisition time point to generate a total memory occupancy sequence, calculate the ratio of the memory occupancy sequence to the total memory occupancy sequence based on the same acquisition time point, and summarize all the ratios based on the acquisition time point to generate a ratio sequence, wherein all objects refer to various data units or data structure instances processed by operators in any group Group in the analysis software; Building a prediction model based on a neural network, wherein the prediction model predicts the memory occupancy sequence based on the proportion sequence and then outputs a first prediction result; The group memory manager checks any of the groups Group, and if the first prediction result is greater than a second threshold, determines that it is abnormal and terminates the execution of the operator; A group container is established, and the group container obtains all objects to be counted based on a weak reference mechanism. When all the objects to be counted in the group container are reclaimed by the garbage collector of the JVM, the group Group is deleted based on the group memory manager, and the memory of the group Group is released and allocated to the remaining group Groups for use.
7. The memory resource usage control method of Java data analysis software according to claim 1, characterized in that: The setting of the computing node further includes: The type of the computing node is set based on the hardware device of the analysis software, and a network interface, a storage device and a computing device are configured for the computing node. The supervisory process of the analysis software distributes the application code to all computing nodes participating in the computing operation, obtains the supervisory process of the computing node and generates process instructions. The process of the computing node generates computing operations based on the process instructions, and the computing device executes the computing operations.
8. The memory resource usage control method of Java data analysis software according to claim 7, characterized in that: The allocation of computing operations based on the computing nodes includes: Based on the storage device, a computing node with the smallest computing load is extracted from all the computing nodes and set as a candidate computing node. All parameter information required for the computing operation is sent to the candidate computing node based on the RPC method. The candidate computing node completes the calculation and outputs the calculation result. The calculation result is sent to the current computing node based on the RPC method. This step is repeated until the computing operation is completed.
9. The memory resource usage control method of Java data analysis software according to claim 1, characterized in that: The transmitting the grouping information based on the execution environment and the thread context includes: Creating a corresponding execution environment for each analysis task, the execution environment receiving a processing instruction corresponding to the grouping information, and the analysis software acquiring a data set and an analysis type of the grouping information based on the processing instruction; Based on the thread manager, multiple threads are set, the data set and the analysis type are passed to the thread context based on the thread parameters, the multiple threads are executed in parallel, each thread processes one data set, and the thread results are output based on the analysis type, and the thread results of all threads are aggregated to generate the analysis result.
10. A memory resource usage control system for Java data analysis software, characterized in that: The memory resource usage control system of the Java data analysis software includes: A logical grouping module is used to logically group the memory resources of the analysis software, generate multiple groups, manage the memory occupancy of the group based on the group memory manager, set the memory corresponding to the group, the analysis software obtains the group information of the group after receiving the request, transmits the group information based on the execution environment and thread context, and maintains the memory of the group based on the group memory manager, wherein the group represents the basic unit of resource application; A first statistical module is used to define a column storage structure corresponding to the group Group for different data types, the column storage structure obtains a current memory size based on a defined variable, and counts a first memory occupancy of the column storage structure based on the memory size, wherein the first memory occupancy refers to the memory space occupied by the column storage structure itself, excluding the memory temporarily occupied during data processing; A second statistical module is used to count the second memory occupancy of any operator during the calculation process, and to combine the second memory occupancy based on the number of types of the operators to generate the computing memory used by the analysis software, wherein the second memory occupancy refers to the memory space temporarily occupied by the operator during the calculation process; a detection module, configured to combine the first memory occupancy with the computing memory to set the combined memory occupancy, set a sampling frequency, collect the combined memory occupancy of the analysis software when it is running based on the sampling frequency to generate a memory occupancy sequence, and the group memory manager performs memory overflow detection based on the memory occupancy sequence; The resource isolation module sets the associated calculation as a calculation operation, sets the calculation node, and allocates the calculation operation based on the calculation node to achieve resource isolation if the calculation task of the analysis software includes associated calculation and the data volume of the associated calculation is greater than a first threshold.
Citation Information
Patent Citations
Memory calculation method and device of java object and electronic equipment
CN118394602A
Method and device for monitoring use condition of Java program memory
CN112835765A
Memory monitoring method and device of Spark platform and storage medium
CN116340094A