A Method for Mutual Invocation between Java and Python in the Form of a Gateway
By establishing a multi-dimensional data acquisition and analysis unit between Java and Python, calculating multiple coefficient values and optimizing the calling strategy through dynamic monitoring and real-time early warning mechanisms, the data omissions and analysis results deviations in the data acquisition and analysis process of Java and Python in the existing technology are solved, and the quality level of cross-language data processing and analysis is significantly improved.
Patent Information
- Application Number
- CN202411759575.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-12-03
AI Technical Summary
The existing gateway forms Java and Python call methods that are mutually called by each other have problems such as data omissions, missing key data fields and biased analysis results in the process of data collection and data analysis, resulting in the weakening of the comprehensiveness and reliability of cross-language data analysis.
By establishing a multi-dimensional data acquisition and analysis unit between Java and Python, including communication performance data acquisition and analysis, functional execution data acquisition and analysis, and data processing and comprehensive evaluation steps, multiple coefficient values are calculated to evaluate the overall effect of the calling process, and optimize the calling strategy through dynamic monitoring and real-time early warning mechanisms.
It significantly improves the accuracy and completeness of data collection and analysis of Java and Python calls, improves the quality of cross-language data processing and analysis, and ensures data support capabilities for intelligent decision-making and efficient operations.
Smart Images

Figure CN119718479B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mutual calling between Java and Python, and more specifically, to a method for mutual calling between Java and Python in a gateway form. Background Art
[0002] With the booming development of the big data era and the continuous deepening of enterprise digital transformation, in many complex business scenarios, the mutual calling method between Java and Python in the form of a gateway has become the core link for integrating multi-source data and mining deep value, and the reliance on its calling accuracy and data integrity is increasing day by day. Relying on this calling method to achieve cross-language data collaborative processing can not only give full play to the advantages of Java's robust architecture and Python's flexible data mining expertise, but also unlock new business insights, bringing unprecedented opportunities for enterprise intelligent decision-making and efficient operations, but also inevitably facing a series of harsh tests. Accurate and complete data interaction and analysis, like a cornerstone, supports the solid building of cross-language collaborative operations, and is the key to meeting the urgent needs of the business for refined and scientific operations.
[0003] The existing gateway-based method for mutual calling between Java and Python mainly includes data collection steps, data transmission steps, data analysis steps, and result integration steps. In the data collection step, collection components that adapt to multiple data sources are deployed on the Java side and the Python side respectively. For diversified data sources such as database records, log files, and real-time sensor data streams, customized collection scripts and adapter interfaces are used to collect various types of data information, and to lay a solid data foundation for subsequent in-depth analysis. In the data transmission step, with the help of efficient and reliable network transmission protocols, such as the internal communication protocol optimized and customized based on TCP / IP, and with advanced data serialization and deserialization mechanisms, the data collected from both ends are standardized and packaged and quickly transmitted to ensure the lossless and timely flow of data in a cross-language environment, and to enhance the stability and integrity of data during transmission. In the data analysis step, according to business goals, Java's high-performance computing library and Python's rich data science toolkit are called, and machine learning models, statistical analysis algorithms, etc. are used to conduct in-depth analysis of the transmitted and aggregated data, so as to use their strengths to explore the potential value of data and provide quantitative support for decision-making. Using the result integration step, the results analyzed by different language terminals are summarized and sorted according to a unified data format and business logic framework, which is convenient for business personnel to clearly interpret and intuitively apply, and achieve synergy and efficiency of cross-language analysis results.
[0004] However, in the actual application process, this set of methods is still troubled by many thorny problems. In the data collection link, due to the complexity and diversity of the interface standards of different data sources and the insufficient adaptability of the two-end collection components to some niche data sources, data omission and the lack of some key data fields often occur. Just like losing key pieces of a jigsaw puzzle, the integrity of the collected data is greatly reduced, seriously weakening the comprehensiveness and reliability of subsequent data analysis; in the data analysis stage, due to the complexity of the coordination between the built-in analysis tools of Java and Python in dealing with complex data structures and heterogeneous data features, it is easy to fall into the local optimum trap. Coupled with the inaccurate control of the subtle differences in cross-language data type conversion, the analysis results are deviated and the insights are not sharp enough, unable to accurately locate business pain points and capture potential opportunities, greatly reducing the practical value and decision-making guidance effectiveness of cross-language data analysis.
[0005] Therefore, there is an urgent need for a method of mutual call between Java and Python in the form of a gateway to fully resolve the existing dilemmas of data collection loopholes and poor accuracy of analysis results, further improve the quality level of cross-language data processing and analysis, and build a solid "data foundation" for intelligent decision-making and efficient operation in the enterprise's digital transformation journey, helping the business to ride the waves and move forward steadily in the fierce market competition. Summary of the Invention
[0006] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a method for mutual call between Java and Python in the form of a gateway, through the following solutions to solve the problems raised in the above-mentioned background technology.
[0007] To achieve the above object, the present invention provides the following technical solution: A method for mutual call between Java and Python in the form of a gateway, including:
[0008] S1. Data batch division step: used to determine the data to be collected as target data, divide the target data into different batches in the way of equal-time division, and mark them as 1, 2,..., n in sequence;
[0009] S2. Data collection step: including a communication performance data collection unit and a function execution data collection unit, used to collect the target data in real time and transmit the collected data to the data processing step; the communication performance data collection unit is used to collect connection establishment data, data transmission data, and connection closing data; the function execution data collection unit is used to collect function call data, data processing data, and task completion data;
[0010] S3. Data processing step: It includes a communication performance data analysis unit and a function execution data analysis unit, which are used to analyze the data transmitted in the data collection step and transmit the analysis results to the data comprehensive evaluation step; the communication performance data analysis unit includes a connection establishment data analysis node, a data transmission data analysis node, and a connection closure data analysis node; the function execution data analysis unit includes a function call data analysis node, a data processing data analysis node, and a task completion data analysis node;
[0011] S4. Data comprehensive evaluation step: It includes a Java and Python mutual call data analysis unit, which is used to comprehensively evaluate the analysis results transmitted in the data processing step and transmit the evaluation results to the real-time warning feedback step;
[0012] S5. Real-time warning feedback step: It is used to preset the comprehensive evaluation index value of Java and Python mutual calls, judge the comprehensive evaluation index value of Java and Python mutual calls according to the preset comprehensive evaluation index value of Java and Python mutual calls, and send corresponding signals according to the judgment results.
[0013] Preferably, the connection establishment data includes the cold start handshake success rate Ft, the network switch reconnection time Ra, and the concurrent connection competition winning rate Cr; the data transmission data includes the large file chunk transmission rate Br, the data packet loss retransmission delay Pr, and the multi-protocol mixed transmission efficiency ratio Td; the connection closure data includes the normal closure resource rollback duration Gt, the abnormal closure data retention integrity Ge, and the idle connection automatic cleaning efficiency Ec.
[0014] Preferably, the function call data includes the cross-language function warm-up time Fw, the concurrent function call conflict rate Fc, and the function call depth backtracking success rate Fs; the data processing data includes the heterogeneous data format adaptation accuracy Da, the data cache aging accuracy Ca, and the data processing pipeline balance Pb; the task completion data includes the task interruption recovery success rate Rt and the resource peak redundancy Rp.
[0015] Preferably, the connection establishment data analysis node is used to establish a connection establishment data calculation model, import the connection establishment data transmitted in the data collection step into the connection establishment data calculation model, and obtain the connection agility adaptation coefficient value. The connection establishment data calculation model is specifically expressed as:
[0016] ,
[0017] where α i represents the connection agility adaptation coefficient value of the i-th calculation, Ft i represents the cold start handshake success rate of the i-th collection, Ra iRepresents the network handover and reconnection time-consuming for the i-th collection, Cr i Represents the concurrent connection competition winning rate for the i-th collection.
[0018] Preferably, the data transmission data analysis node is used to establish a data transmission data calculation model, import the data transmission data transmitted in the data collection step into the data transmission data calculation model, and obtain a transmission efficient cooperation coefficient value. The data transmission data calculation model is specifically expressed as:
[0019] ,
[0020] where β i Represents the transmission efficient cooperation coefficient value for the i-th calculation, Br i Represents the large file chunk transmission rate for the i-th collection, Pr i Represents the data packet loss retransmission delay for the i-th collection, Td i Represents the multi-protocol mixed transmission efficiency ratio for the i-th collection.
[0021] Preferably, the connection closing data analysis node is used to establish a connection closing data calculation model, import the connection closing data transmitted in the data collection step into the connection closing data calculation model, and obtain a connection health closed-loop coefficient value. The connection closing data calculation model is specifically expressed as:
[0022] ,
[0023] where γ i Represents the connection health closed-loop coefficient value for the i-th calculation, Gt i Represents the normal closing resource rollback duration for the i-th collection, Ge i Represents the integrity of abnormal closing data retention for the i-th collection, Ec i Represents the automatic cleaning efficiency of idle connections for the i-th collection.
[0024] Preferably, the function call data analysis node is used to establish a function call data calculation model, import the function call data transmitted in the data collection step into the function call data calculation model, and obtain a function call robustness coefficient value. The function call data calculation model is specifically expressed as:
[0025] ,
[0026] where ∂ i Represents the function call robustness coefficient value for the i-th calculation, Fw i Represents the cross-language function warm-up time-consuming for the i-th collection, Fc i Represents the concurrent function call conflict rate for the i-th collection, Fs iRepresents the success rate of function call depth backtracking for the i-th collection.
[0027] Preferably, the data processing data analysis node is used to establish a data processing data calculation model, import the data processing data transmitted in the data collection step into the data processing data calculation model, and obtain a data processing excellence coefficient value. The data processing data calculation model is specifically expressed as:
[0028] ,
[0029] where ℓ i represents the data processing excellence coefficient value calculated for the i-th time, Da i represents the accuracy of heterogeneous data format adaptation collected for the i-th time, Ca i represents the accuracy of data cache aging collected for the i-th time, Pb i represents the balance degree of the data processing pipeline collected for the i-th time.
[0030] Preferably, the task completion data analysis node is used to establish a task completion data calculation model, import the task completion data transmitted in the data collection step into the task completion data calculation model, and obtain a task completion satisfaction coefficient value. The task completion data calculation model is specifically expressed as:
[0031] ,
[0032] where η i represents the task completion satisfaction coefficient value calculated for the i-th time, Rt i represents the success rate of task interruption recovery collected for the i-th time, Rp i represents the resource peak redundancy collected for the i-th time.
[0033] Preferably, the Java and Python mutual call data analysis unit is used to establish a Java and Python mutual call data calculation model, import the connection agility adaptation coefficient value, transmission efficiency coordination coefficient value, connection health closed-loop coefficient value, function call robustness coefficient value, data processing excellence coefficient value, and task completion satisfaction coefficient value transmitted in the data processing step into the Java and Python mutual call data calculation model, and obtain a Java and Python mutual call comprehensive evaluation index value. The Java and Python mutual call data calculation model is specifically expressed as:
[0034] ,
[0035] where IA represents the calculated Java and Python mutual call comprehensive evaluation index value, α i represents the connection agility adaptation coefficient value calculated for the i-th time, β iRepresents the value of the transmission efficient collaboration coefficient for the i-th calculation, γ i Represents the value of the connection health closed-loop coefficient for the i-th calculation, ∂ i Represents the value of the function call robustness coefficient for the i-th calculation, ℓ i Represents the value of the data processing excellence coefficient for the i-th calculation, η i Represents the value of the task completion satisfaction coefficient for the i-th calculation.
[0036] Technical effects and advantages of the present invention:
[0037] Through the data batch division step, the present invention effectively divides the data and operation processes of mutual calls between Java and Python, greatly improving the pertinence and effectiveness of cross-language call data. By collecting connection establishment data, data transmission data, connection closing data, function call data, data processing data, and task completion data in multiple dimensions, this multi-faceted data collection mode effectively breaks through the limitation of incomplete data docking when traditional Java and Python call each other, providing a rich and reliable information basis for subsequent accurate and smooth mutual calls;
[0038] Through data parsing, the present invention deeply analyzes the data collected from both Java and Python ends, and then calculates the connection agility adaptation coefficient value, transmission efficient collaboration coefficient value, connection health closed-loop coefficient value, function call robustness coefficient value, data processing excellence coefficient value, and task completion satisfaction coefficient value, clearly pointing out the factors that may affect the mutual call effect between Java and Python; By comprehensively considering and organically integrating the data parsing results, the comprehensive evaluation index value of the entire mutual call process between Java and Python is obtained, significantly improving the scientificity and feasibility of the mutual call scheme;
[0039] Through dynamic monitoring and continuous follow-up, once a call deviation or data interaction anomaly is detected, a warning message is immediately sent. Relevant staff can timely understand the status and potential risk points of the cross-language call system, and quickly take measures to adjust and improve the call strategy and error handling plan, providing a solid guarantee for achieving efficient and stable mutual calls between Java and Python in the form of a gateway and sustainable cross-language collaboration. Brief Description of the Drawings
[0040] Figure 1 Is the overall structural schematic diagram of the present invention. Detailed Embodiments
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] As shown in the attached Figure 1 A method for mutual call between Java and Python in a gateway form, including:
[0043] S1. Data batch division step: used to determine the data to be collected as target data, divide the target data into different batches in an equal-time division manner, and sequentially label them as 1, 2,..., n.
[0044] It should be specifically noted in this embodiment that the equal-time division method is divided according to the characteristics of each target data. For the real-time requirements of each type of target data, it is collected from the database at equal intervals of time and divided into different collection batches.
[0045] S2. Data collection step: including a communication performance data collection unit and a function execution data collection unit, used to collect target data in real time and transmit the collected data to the data processing step; the communication performance data collection unit is used to collect connection establishment data, data transmission data, and connection closing data; the function execution data collection unit is used to collect function call data, data processing data, and task completion data.
[0046] In this embodiment, it should be specifically noted that the connection establishment data includes the cold start handshake success rate Ft, the network switch reconnection time Ra, and the concurrent connection competition winning rate Cr; the data transmission data includes the large file chunk transmission rate Br, the data packet loss retransmission delay Pr, and the multi-protocol mixed transmission efficiency ratio Td; the connection closing data includes the normal closing resource rollback duration Gt, the abnormal closing data retention integrity Ge, and the idle connection automatic cleaning efficiency Ec; the function call data includes the cross-language function warm-up time Fw, the concurrent function call conflict rate Fc, and the function call depth backtracking success rate Fs; the data processing data includes the heterogeneous data format adaptation accuracy Da, the data cache aging accuracy Ca, and the data processing pipeline balance Pb; the task completion data includes the task interruption recovery success rate Rt and the resource peak redundancy Rp.
[0047] The high success rate of the cold start handshake proves that the system can quickly adapt to the initial network environment, allocate its own resources, and accurately build a communication bridge according to the established protocol (similar to the TCP three-way handshake logic, Java sends a SYN request, Python responds with SYN+ACK, and Java then returns ACK), avoiding delays in starting services. It is of great significance for scenarios that require timely start of interactions such as real-time monitoring and scheduled task pushing, and is a key indicator to measure the initial communication reliability of the system. Its collection method is as follows: In the underlying code of the communication framework between the Java side and the Python side, a handshake status recording module is embedded. Each time the system performs a cold start (which can be judged based on the time interval between the system startup timestamp and the end time of the last communication. If the interval exceeds 1 hour, it is considered a cold start), when initiating a connection, mark "handshake attempt starts" and record the time t1. When the confirmation from the other party is received completely (the three-way handshake is completed), mark "handshake successful" and record t2. At the same time, count the total number of cold starts totalColdStarts and the number of successful cold starts successColdStarts, and calculate the cold start handshake success rate according to the formula cold start handshake success rate = (successColdStarts / totalColdStarts) * 100%.
[0048] The time taken for network switch reconnection intuitively reflects the system's adaptability. A short reconnection time can reduce service lags and the sense of data interruption, ensuring uninterrupted video conference venue switching, off-site data synchronization, etc., maintaining the system's "always online" status, and improving user experience and service continuity. Its collection method is as follows: Use the system network status monitoring interface (Java can use the NetworkInterface class and related libraries to monitor network changes, and Python uses the psutil library to monitor the network card status). When a network switch is detected (such as IP address change, network interface device switch), record the startTime at the moment when the reconnection logic is triggered. When the reconnection is successful (after the Java-side Socket reconnects, it can send and receive test data normally, and Python is similar to confirm according to the feedback of the socket module), record the endTime. The difference endTime - startTime is the time taken for network switch reconnection. Calculate the average value of multiple rounds of switch records and count the maximum and minimum values for auxiliary analysis.
[0049] The concurrent connection competition winning rate determines whether the system can quickly accept numerous clients. A high winning rate means that the system efficiently allocates resources such as ports, memory, and threads, avoids connection blocking and queuing, and ensures stable operation when a large amount of traffic surges instantaneously. It is a "pioneer indicator" for the anti-pressure of large distributed systems. The acquisition method is as follows: Write a concurrent connection test tool (using an ExecutorService thread pool in Java to batch create connection tasks, and using concurrent.futures.ThreadPoolExecutor in Python) to simulate concurrent requests from multiple clients. Each task carries an independent identifier and a status flag (initially "in competition"). When a connection is successfully established, it is updated to "won", and when it fails, it is recorded as "failed". Count the total number of concurrent connections totalConcurrent and the number of successful connections successConcurrent, and calculate the concurrent connection competition winning rate according to the formula: (successConcurrent / totalConcurrent) * 100%.
[0050] The large file chunk transfer rate is a "touchstone" for the system's transmission capacity. A fast chunk transfer rate ensures efficient file transfer, reduces transmission time, alleviates long-term high loads on the gateway and network, and is conducive to timely data synchronization and release of local storage. The acquisition method is as follows: Before transmitting a large file on the Java side (reading with FileInputStream), split it into chunks according to a set chunk size (such as 1MB) and record the start time startTime of the first chunk transmission. When the corresponding chunk is received on the Python side, record the end time endTime. Count the number of transmitted chunks blockCount and the total time consumption. Calculate the number of transmitted bytes based on the file chunk size (1MB per chunk), and calculate the large file chunk transfer rate according to the formula: (number of transmitted bytes / (endTime - startTime)) / 1024 / 1024 (converted to MB / s).
[0051] A short data packet loss retransmission delay can quickly restore the data sequence, ensure that the receiving end processes data in order, respond promptly to data changes, avoid incorrect decisions caused by packet loss, and maintain accurate operation of the service. The acquisition method is as follows: The sending end (Java or Python) assigns a unique sequence number (an incrementing integer) to each real-time data packet. The receiving end maintains a list of expected received sequence numbers. When a packet loss is detected (the received sequence numbers are not consecutive), record the packet loss time lostTime. When the retransmitted packet arrives, record the recoverTime. recoverTime - lostTime is the data packet loss retransmission delay.
[0052] The multi - protocol mixed - transmission efficiency ratio can measure the system's ability to coordinate multiple protocols, allocate bandwidth reasonably, and avoid protocol conflicts. A higher ratio indicates better protocol coordination. For example, in an intelligent office platform where both document transmission and instant messaging operate smoothly, the acquisition method is as follows: at the gateway traffic monitoring point (which can be a hardware packet - capture device or a software proxy), the total throughput (totalThroughput, in bytes per second) during multi - protocol mixed - transmission is separately counted. Then, the throughput in the optimal transmission scenario for each protocol (optimalThroughput_i corresponds to the i - th protocol) is tested individually. The multi - protocol mixed - transmission efficiency ratio is calculated according to the formula: multi - protocol mixed - transmission efficiency ratio=(totalThroughput / sum(optimalThroughput_i))*100%.
[0053] The duration of normal resource rollback when closing affects subsequent connection creation and the overall system performance. A short resource rollback duration ensures that the system can start afresh and continuously and efficiently handle new service requests, which is particularly important in high - concurrency short - connection scenarios (such as frequent interactions in microservices). The acquisition method is as follows: at the start of the connection - closing code logic (before the combined operations of Java socket.close(), Python socket.shutdown() and socket.close()), use System.currentTimeMillis() (in Java) or time.time()*1000 (in Python, converted to milliseconds) to record startTime. Monitor the moment when the resources are completely released with the help of operating - system resource monitoring tools (lsof in Linux to check file handles, netstat to view ports, Windows Task Manager and corresponding network commands) and record it as endTime. endTime - startTime is the duration of normal resource rollback when closing.
[0054] The integrity of data retention during abnormal closure is for subsequent business recovery and fault troubleshooting. It is a "safety net" to protect data assets and reduce losses, and is indispensable for data - sensitive services to prevent key data loss and maintain business continuity. The acquisition method is as follows: within the connection abnormal - closure capture logic (Java try - catch to capture exceptions such as IOException, Python try - except to catch socket.error, etc.), record information such as the location and size of the unprocessed data that has been transmitted. Compare with the total data volume before the abnormality, count the total number of abnormal closures totalAbnormal and the number of times of successful data retention successRetain. The integrity of data retention during abnormal closure is calculated according to the formula: integrity of data retention during abnormal closure=(successRetain / totalAbnormal)*100%.
[0055] The idle connection automatic cleanup time limit means that timely cleanup of long-term idle connections that consume system resources can optimize resource allocation, improve the system's response speed to new connections, and enhance overall operating efficiency. For example, cleaning idle connections during the low-load period at night on the server saves energy and improves efficiency. The collection method is: the system runs an idle connection monitoring thread (Java timer task, Pythonthreading.Timer implementation) in the background, scans the connection pool according to the set idle threshold (such as 5 minutes of no data interaction), and starts the cleanup timer when an idle connection is found. The start of the cleanup startClean and the end of the cleanup endClean are recorded (based on the completion of resource release or connection mark deletion). endClean-startClean is the idle connection automatic cleanup time limit.
[0056] The cross-language function warm-up time determines the "initial speed" of the function response. Short time can quickly activate the function and put it into business operation, reducing user waiting. For example, in scenarios such as new function launch and low-frequency business function call, the system agility is improved and the "frustration" of cold start is avoided. The collection method is: before Java calls the Python function (or vice versa) entry code, use System.currentTimeMillis() (Java) or time.time()*1000 (Python) to record startTime. Function calls involve loading dependent libraries (Pythonimport modules, JavaClassLoader loading classes), initializing the environment (such as Python interpreter configuration, Java static variable initialization), and recording endTime until the function can execute the first instruction. EndTime-startTime is the cross-language function warm-up time.
[0057] The low conflict rate of concurrent function calls can ensure that each call is executed independently and correctly, avoid data confusion and result errors, and is the key to ensuring the stability and accuracy of large-scale parallel processing of the system. The collection method is: build a concurrent test framework (JavaExecutorService multi-threaded concurrency, Pythonconcurrent.futures.ThreadPoolExecutor) to initiate batch function calls, and associate each call with a unique identifier and status (initially "running"). Errors caused by contention for shared resources (global variables, database connection pool limited connections) are updated to "conflict", and the total concurrent number of totalConcurrentCall and the number of conflicts conflictCount are counted, and the concurrent function call conflict rate = (conflictCount / totalConcurrentCall)*100% is calculated.
[0058] The success rate of function call depth backtracking measures the frequency of whether a recursive call can successfully backtrack to the upper-level call after reaching a certain depth and continue to execute normally. A high success rate means that the program can safely and effectively resume execution after the recursive call ends, which is crucial for the stability and efficiency of the program. It is the "compass" for system operation and maintenance, improving the efficiency of fault troubleshooting, shortening the repair cycle, ensuring the continuous and healthy operation of the business, reducing operation and maintenance costs, especially prominent when complex business logic functions go wrong. Its collection method is as follows: Embed a log record and stack frame capture mechanism at the bottom layer of the function call framework (Java uses the StackTraceElement array to record the call stack, and the Python traceback module captures stack information). Each time the function finishes execution or exits abnormally, parse the log and stack frame to locate the root cause of the error, count the total number of function calls totalCall and the number of successful backtracking repairs successBacktrace, and calculate according to the formula: success rate of function call depth backtracking = (successBacktrace / totalCall) * 100%.
[0059] The accuracy rate of heterogeneous data format adaptation ensures error-free "translation" of data across languages, maintains the integrity of data semantics and structure, and avoids "misreading" of the business due to data mismatch. It lays a solid foundation for businesses with frequent data fusion and interaction (cross-language collaboration in data analysis pipelines). Its collection method is as follows: In the data conversion module (Java custom conversion class, Python conversion function group), for each format conversion operation, clone a data copy before conversion (Java clone() or serialization / deserialization, Python copy.deepcopy()), and compare the copy with the target data element by element and field by field after conversion (according to data type rules, compare numerical values for numbers and compare contents for strings, etc.). Count the total number of conversions totalConversion and the number of accurate conversions successConversion, and calculate according to the formula: accuracy rate of heterogeneous data format adaptation = (successConversion / totalConversion) * 100%.
[0060] The accuracy of data cache aging refers to the ability of the data aging (eviction) strategy in the cache system to accurately reflect the access frequency and timeliness of data, so as to ensure that the most valuable data is retained in the cache. High accuracy ensures the retention of hot data and the elimination of cold data, maintains the healthy "metabolism" of the cache, improves the cache hit rate, releases space, accelerates data acquisition. The acquisition method is as follows: in the cache management component (Java can be extended based on GuavaCache, Python uses functools.lru_cache and custom cleaning logic), for each cleaning operation (triggered by a timer or when the cache is full), traverse the cleaning data list, compare the actually cleaned data (expired, low-frequency) with the data that should be cleaned, and count the total number of cleaning operations totalClean and the number of accurate cleaning operations successClean. Calculate the accuracy of data cache aging according to the formula: (successClean / totalClean) * 100%.
[0061] It refers to whether the processing capabilities and loads at each stage in the data processing pipeline are balanced to ensure the efficient operation of the entire pipeline. High balance avoids bottlenecks, realizes efficient resource utilization and smooth processes, improves overall throughput, shortens the cycle from raw data to finished products, and enhances the efficiency in big data batch processing and real-time stream processing. The acquisition method is as follows: embed timestamp records (Java: System.currentTimeMillis(), Python: time.time() * 1000) at the start and end of each link in the pipeline, count the processing time processTime_i (i represents the link), calculate the longest link time maxTime, and calculate the balance of the data processing pipeline according to the formula: (sum(processTime_i) / (len(processTime_i) * maxTime)) * 100%.
[0062] The success rate of task interruption recovery is the "touchstone" of system resilience. A high success rate reduces repetitive labor and data loss, ensures business continuity, and resists risks and maintains operation in critical businesses (such as financial transactions, scientific research operations). The acquisition method is as follows: simulate interruption scenarios (actively shut down processes, cut off the network to simulate failures). After the task recovery mechanism (based on logs, checkpoint to restore progress) is started, count the total number of task interruptions totalInterrupt and the number of successful recoveries successRecover, and calculate the success rate of task interruption recovery according to the formula: (successRecover / totalInterrupt) * 100%.
[0063] A high peak redundancy of task resources implies resource waste and poor configuration. Optimizing this metric can rationally allocate resources, enhance the system's concurrent multi-tasking capacity, avoid resource contention, ensure that each task "has enough but not too much", and operate smoothly. The acquisition method is as follows: With the help of system resource monitoring tools (such as Linuxtop and htop for CPU and memory, and Windows Task Manager), record the resource occupancy (CPU usage rate, memory amount) at each stage of task execution, estimate the optimal resource amount optimalResource based on task complexity and historical average, count the peak resource amount peakResource, and calculate the task resource peak redundancy according to the formula: task resource peak redundancy = ((peakResource - optimalResource) / optimalResource)*100%.
[0064] S3. Data processing steps: Include a communication performance data analysis unit and a function execution data analysis unit, which are used to analyze the data transmitted in the data acquisition step and transmit the analysis results to the data comprehensive evaluation step; the communication performance data analysis unit includes a connection establishment data analysis node, a data transmission data analysis node, and a connection closure data analysis node; the function execution data analysis unit includes a function call data analysis node, a data processing data analysis node, and a task completion data analysis node.
[0065] In this embodiment, specifically, it should be noted that: the connection establishment data analysis node is used to establish a connection establishment data calculation model, import the connection establishment data transmitted in the data acquisition step into the connection establishment data calculation model, and obtain the connection agility adaptation coefficient value. The connection establishment data calculation model is specifically expressed as:
[0066] ,
[0067] where α i represents the connection agility adaptation coefficient value of the i-th calculation, Ft i represents the cold start handshake success rate of the i-th acquisition, Ra i represents the network handover reconnection time of the i-th acquisition, Cr i represents the concurrent connection competition winning rate of the i-th acquisition.
[0068] In this embodiment, specifically, it should be noted that: the data transmission data analysis node is used to establish a data transmission data calculation model, import the data transmission data transmitted in the data acquisition step into the data transmission data calculation model, and obtain the transmission efficient collaboration coefficient value. The data transmission data calculation model is specifically expressed as:
[0069] ,
[0070] where βi Represents the value of the transmission efficient collaboration coefficient for the i-th calculation, Br i Represents the transmission rate of large file chunks collected for the i-th time, Pr i Represents the data packet loss retransmission delay for the data collected for the i-th time, Td i Represents the multi-protocol mixed transmission efficiency ratio for the data collected for the i-th time.
[0071] In this embodiment, it should be specifically noted that: the connection close data analysis node is used to establish a connection close data calculation model, import the connection close data transmitted in the data collection step into the connection close data calculation model, and obtain the connection health closed-loop coefficient value. The connection close data calculation model is specifically expressed as:
[0072] ,
[0073] Among them, γ i Represents the value of the connection health closed-loop coefficient for the i-th calculation, Gt i Represents the normal close resource rollback duration collected for the i-th time, Ge i Represents the integrity of the abnormal close data retention collected for the i-th time, Ec i Represents the automatic cleaning time limit of idle connections collected for the i-th time.
[0074] In this embodiment, it should be specifically noted that: the function call data analysis node is used to establish a function call data calculation model, import the function call data transmitted in the data collection step into the function call data calculation model, and obtain the function call robustness coefficient value. The function call data calculation model is specifically expressed as:
[0075] ,
[0076] Among them, ∂ i Represents the value of the function call robustness coefficient for the i-th calculation, Fw i Represents the cross-language function warm-up time consumed for the i-th time, Fc i Represents the concurrent function call conflict rate collected for the i-th time, Fs i Represents the success rate of function call depth backtracking collected for the i-th time.
[0077] In this embodiment, it should be specifically noted that: the data processing data analysis node is used to establish a data processing data calculation model, import the data processing data transmitted in the data collection step into the data processing data calculation model, and obtain the data processing excellence coefficient value. The data processing data calculation model is specifically expressed as:
[0078] ,
[0079] Among them, ℓ iRepresents the data processing excellence coefficient value for the i-th calculation, Da i Represents the accuracy of heterogeneous data format adaptation for the i-th collection, Ca i Represents the accuracy of data cache aging for the i-th collection, Pb i Represents the balance of the data processing pipeline for the i-th collection.
[0080] In this embodiment, specifically, it should be noted that: the task completion data analysis node is used to establish a task completion data calculation model, import the task completion data transmitted in the data collection step into the task completion data calculation model, and obtain the task completion satisfaction coefficient value. The task completion data calculation model is specifically expressed as:
[0081] ,
[0082] Among them, η i Represents the task completion satisfaction coefficient value for the i-th calculation, Rt i Represents the success rate of task interruption recovery for the i-th collection, Rp i Represents the resource peak redundancy for the i-th collection.
[0083] S4. Data comprehensive evaluation step: It includes a Java and Python mutual call data analysis unit, which is used to comprehensively evaluate the analysis results transmitted in the data processing step and transmit the evaluation results to the real-time warning feedback step.
[0084] In this embodiment, specifically, it should be noted that: the Java and Python mutual call data analysis unit is used to establish a Java and Python mutual call data calculation model, import the connection agility adaptation coefficient value, transmission efficiency coordination coefficient value, connection health closed-loop coefficient value, function call robustness coefficient value, data processing excellence coefficient value, and task completion satisfaction coefficient value transmitted in the data processing step into the Java and Python mutual call data calculation model, and obtain the Java and Python mutual call comprehensive evaluation index value. The Java and Python mutual call data calculation model is specifically expressed as:
[0085] ,
[0086] Among them, IA represents the calculated Java and Python mutual call comprehensive evaluation index value, α i Represents the connection agility adaptation coefficient value for the i-th calculation, β i Represents the transmission efficiency coordination coefficient value for the i-th calculation, γ i Represents the connection health closed-loop coefficient value for the i-th calculation, ∂ i Represents the function call robustness coefficient value for the i-th calculation, ℓ iRepresents the data processing excellence coefficient value of the i-th calculation, η i Represents the task completion satisfaction coefficient value of the i-th calculation.
[0087] S5. Real-time warning feedback step: Used for the preset value of the comprehensive evaluation index of Java and Python mutual calls. Judge the comprehensive evaluation index value of Java and Python mutual calls according to the preset value of the comprehensive evaluation index of Java and Python mutual calls, and send corresponding signals according to the judgment results.
[0088] In this embodiment, specifically, it should be noted that: the preset value of the comprehensive evaluation index of Java and Python mutual calls is marked as IA def , when IA def <= IA, a normal signal is sent. This signal indicates that the preset value of the comprehensive evaluation index of Java and Python mutual calls is less than or equal to the comprehensive evaluation index value of Java and Python mutual calls, indicating that the mutual call situation between Java and Python is good; and the following operations are performed:
[0089] SS1. Network connection: Realize communication by establishing a network connection between Python and JVM. This connection is based on the TCP / IP protocol, allowing Python programs to directly call the methods of Java objects without complex serialization and deserialization operations;
[0090] SS2. GatewayServer: On the Java side, use the GatewayServer class to start a server. This server allows Python programs to communicate with the JVM through local network sockets. The GatewayServer instantiates an entity pointer (entrypoint), which can be any object, such as Façade, singleton, list, etc., allowing Python programs to access pre-configured Java objects;
[0091] SS3. JavaGateway: On the Python side, use the JavaGateway class to connect to the JVM. This gateway is responsible for establishing and maintaining the connection with the Java side and provides an interface to access Java objects and methods;
[0092] SS4. Object reference management: Each time a Java object is sent to the Python side, the Gateway class on the Java side will maintain a reference to the object. Once the object on the Python side is garbage collected (reference count is 0), the reference on the Java side will also be removed. If the gateway is closed, the remaining references will also be removed from the Java side;
[0093] SS5. Two-way communication: It not only allows Python to call Java objects, but also supports Java programs to call back Python objects to achieve two-way communication;
[0094] SS6. Memory model: When dealing with references of Java objects on the Python side, the problem of circular references will be involved. JavaObject and JavaMember will reference each other. These objects will not be immediately garbage collected after the last reference on the Python side is removed, but they will eventually be garbage collected provided that Python's garbage collector runs before the program exits.
[0095] When IA def > IA, a warning signal is sent to relevant technical management personnel. This signal indicates that the preset value of the comprehensive evaluation index for the mutual call between Java and Python is greater than the value of the comprehensive evaluation index for the mutual call between Java and Python, indicating that the mutual call situation between Java and Python is poor and relevant technical personnel need to make adjustments.
[0096] Through the data batch division step, the present invention effectively divides the data and operation processes of mutual calls between Java and Python, greatly improving the pertinence and effectiveness of cross-language call data. By collecting data on connection establishment, data transmission, connection closure, function calls, data processing, and task completion from multiple dimensions, this multi-faceted data collection mode effectively breaks through the limitation of incomplete data docking when traditional Java and Python call each other, providing a rich and reliable information basis for subsequent accurate and smooth mutual calls; through data parsing, the data collected from both Java and Python ends is deeply analyzed, and then the connection agility adaptation coefficient value, transmission efficiency coordination coefficient value, connection health closed-loop coefficient value, function call robustness coefficient value, data processing excellence coefficient value, and task completion satisfaction coefficient value are calculated, clearly indicating the factors that may affect the mutual call effect between Java and Python; through comprehensive consideration and organic integration of the data parsing results, the comprehensive evaluation index value of the entire mutual call process between Java and Python is obtained, significantly improving the scientificity and feasibility of the mutual call solution. Through continuous dynamic monitoring, once a call deviation or data interaction anomaly is detected, a warning message is immediately sent, enabling relevant staff to promptly understand the status and potential risk points of the cross-language call system, and quickly take measures to adjust and improve the call strategy and error handling plan, providing a solid guarantee for achieving efficient and stable Java and Python mutual calls based on the gateway form and sustainable cross-language collaboration; dynamic access to Java objects: allows Python programs to dynamically access Java objects in the Java Virtual Machine (JVM), enabling Python programs to call Java methods just like accessing local objects. This dynamic access capability greatly improves development efficiency and flexibility.
[0097] In addition, the present invention also has the following advantages:
[0098] Bidirectional communication: not only supports Python to call Java, but also supports Java to call back Python objects, realizing true bidirectional communication. This bidirectional communication mechanism enables the present invention to perform excellently in dealing with much more complex multi-language integration scenarios;
[0099] High performance: The efficient communication mechanism implemented through network connections avoids complex serialization and deserialization operations, thereby improving performance. According to actual tests, the data transmission speed of the present invention is nearly 50% faster than that of traditional JNI interfaces;
[0100] Easy to integrate: The API of the present invention is designed simply and clearly, making it easy to integrate into existing Python and Java projects. This enables developers to quickly achieve interoperability between Python and Java in the project
[0101] Asynchronous call: The present invention supports the asynchronous call mode. In the asynchronous mode, the Python program can initiate a call request and then immediately continue to execute other tasks without waiting for the response from the Java side. This non-blocking method improves the overall performance of the program.
[0102] Concurrent processing ability: It supports multi-threaded concurrent processing, which means it can handle multiple requests simultaneously. In a high-concurrency environment, the performance advantage of the present invention is particularly obvious.
[0103] Through these technical effects, the present invention is not only powerful in function but also excellent in performance and stability, enabling developers to better utilize the advantages of Python and Java and achieve seamless integration between Python and Java. Generally speaking, the present invention enables Python to seamlessly access and operate Java objects in the JVM by establishing a network connection, and also supports Java's callback to Python to achieve two-way communication. This design not only improves performance but also simplifies the development process.
[0104] Secondly: In the attached drawings of the disclosed embodiments of the present invention, only the structures related to the disclosed embodiments are involved. Other structures can refer to the general design. Without conflict, the same embodiment and different embodiments of the present invention can be combined with each other.
[0105] Finally: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for mutual invocation between Java and Python in the form of a gateway, characterized in that: include: S1, data batch division step: used to determine the data to be collected as target data, divide the target data into different batches in an equal time division manner, and mark them as 1, 2, ..., n in sequence; S2, data collection step: including a communication performance data collection unit and a function execution data collection unit, which are used to collect target data in real time and transmit the collected data to the data processing step; The communication performance data collection unit is used to collect connection establishment data, data transmission data and connection closing data; the function execution data collection unit is used to collect function call data, data processing data and task completion data; S3, data processing step: including a communication performance data analysis unit and a function execution data analysis unit, which are used to analyze the data transmitted in the data collection step and transmit the analysis results to the data comprehensive evaluation step; the communication performance data analysis unit includes a connection establishment data analysis node, a data transmission data analysis node and a connection closing data analysis node; The function execution data analysis unit includes a function call data analysis node, a data processing data analysis node and a task completion data analysis node; S4, data comprehensive evaluation step: including Java and Python calling each other data analysis units, for comprehensive evaluation of the analysis results transmitted in the data processing step, and transmitting the evaluation results to the real-time warning feedback step; S5, real-time warning feedback step: used for the preset value of the comprehensive evaluation index of mutual calls between Java and Python, judging the comprehensive evaluation index value of mutual calls between Java and Python according to the preset value of the comprehensive evaluation index of mutual calls between Java and Python, and sending a corresponding signal according to the judgment result.
2. The method for mutual invocation between Java and Python in the form of a gateway according to claim 1, characterized in that: The connection establishment data includes the cold start handshake success rate Ft, the network switching reconnection time Ra and the concurrent connection competition winning rate Cr; the data transmission data includes the large file block transmission rate Br, the data packet loss retransmission delay Pr and the multi-protocol mixed transmission efficiency ratio Td; the connection closing data includes the normal closing resource rollback time Gt, the abnormal closing data retention integrity Ge and the idle connection automatic cleanup time Ec.
3. The method for mutual invocation between Java and Python in the form of a gateway according to claim 1, characterized in that: The function call data includes the cross-language function warm-up time Fw, the concurrent function call conflict rate Fc and the function call deep backtracing success rate Fs; the data processing data includes the heterogeneous data format adaptation accuracy Da, the data cache aging accuracy Ca and the data processing pipeline balance Pb; the task completion data includes the task interruption recovery success rate Rt and the resource peak redundancy Rp.
4. The method for mutual invocation between Java and Python in the form of a gateway according to claim 1, characterized in that: The connection establishment data analysis node is used to establish a connection establishment data calculation model, import the connection establishment data transmitted in the data collection step into the connection establishment data calculation model, and obtain the connection agility adaptation coefficient value. The connection establishment data calculation model is specifically expressed as: , Among them, α i Indicates the connection agility adaptation coefficient value calculated for the i-th time, Ft i Indicates the cold start handshake success rate of the ith acquisition, Ra i Indicates the network switching and reconnection time of the i-th acquisition, Cr i Indicates the concurrent connection competition winning rate of the i-th collection.
5. The method for mutual invocation between Java and Python in the form of a gateway according to claim 1, characterized in that: The data transmission data analysis node is used to establish a data transmission data calculation model, import the data transmission data transmitted in the data collection step into the data transmission data calculation model, and obtain the transmission high efficiency coordination coefficient value. The data transmission data calculation model is specifically expressed as: , Among them, β i represents the transmission efficiency coordination coefficient value calculated for the i-th time, Br i represents the transmission rate of the large file slices collected for the i-th time, Pr i Indicates the delay in retransmitting lost data packets collected for the i-th time, Td i It represents the multi-protocol mixed transmission efficiency ratio of the i-th acquisition.
6. The method for mutual invocation between Java and Python in the form of a gateway according to claim 1, characterized in that: The connection closing data analysis node is used to establish a connection closing data calculation model, import the connection closing data transmitted in the data collection step into the connection closing data calculation model, and obtain the connection health closed loop coefficient value. The connection closing data calculation model is specifically expressed as: , Among them, γ i Indicates the connection health closed-loop coefficient value calculated for the i-th time, Gt i Indicates the normal closing resource rollback duration of the i-th collection, Ge i Indicates the integrity of abnormal shutdown data collected for the i-th time, Ec i Indicates the automatic cleanup time of idle connections collected for the i-th time.
7. The method for mutual invocation between Java and Python in the form of a gateway according to claim 1, characterized in that: The function call data analysis node is used to establish a function call data calculation model, import the function call data transmitted in the data collection step into the function call data calculation model, and obtain the function call robustness coefficient value. The function call data calculation model is specifically expressed as: , Among them, ∂ i Indicates the robustness coefficient value of the function call calculated for the i-th time, Fw i Fc represents the time taken to warm up the cross-language function collected for the i-th time. i Fs represents the concurrent function call conflict rate of the ith collection. i Indicates the success rate of the function call depth backtracing collected for the i-th time.
8. The method for mutual invocation between Java and Python in the form of a gateway according to claim 1, characterized in that: The data processing data analysis node is used to establish a data processing data calculation model, import the data processing data transmitted in the data collection step into the data processing data calculation model, and obtain the data processing excellence coefficient value. The data processing data calculation model is specifically expressed as: , Among them, ℓ i represents the data processing excellence coefficient value calculated for the i-th time, Da i represents the adaptation accuracy of heterogeneous data formats collected for the i-th time, Ca i represents the aging accuracy of the data cache collected for the i-th time, Pb i Indicates the balance degree of the data processing pipeline for the i-th acquisition.
9. The method for mutual invocation between Java and Python in the form of a gateway according to claim 1, characterized in that: The task completion data analysis node is used to establish a task completion data calculation model, import the task completion data transmitted in the data collection step into the task completion data calculation model, and obtain the task completion degree coefficient value. The task completion data calculation model is specifically expressed as: , Among them, η i Indicates the task completion coefficient value calculated for the i-th time, Rt i Rp represents the success rate of task interruption recovery for the i-th acquisition. i Indicates the resource peak redundancy of the i-th acquisition.
10. The method for mutual invocation between Java and Python in the form of a gateway according to claim 1, characterized in that: The Java and Python mutual call data analysis unit is used to establish a Java and Python mutual call data calculation model, and import the connection agility adaptation coefficient value, transmission efficient coordination coefficient value, connection health closed-loop coefficient value, function call robustness coefficient value, data processing excellence coefficient value and task completion satisfaction coefficient value of the data processing step transmission into the Java and Python mutual call data calculation model to obtain the Java and Python mutual call comprehensive evaluation index value. The Java and Python mutual call data calculation model is specifically expressed as: , Among them, IA represents the calculated comprehensive evaluation index value of Java and Python mutual calls, α i represents the connection agility adaptation coefficient value calculated for the i-th time, β i represents the transmission efficiency coordination coefficient value calculated for the i-th time, γ i represents the connection health closed-loop coefficient value calculated for the i-th time, ∂ i represents the robustness coefficient value of the function call calculated for the i-th time, ℓ i represents the data processing excellence coefficient value calculated for the i-th time, η i Indicates the task completion coefficient value calculated for the i-th time.
Citation Information
Patent Citations
Data processing method and device, data acquisition method and device, equipment and storage medium
CN116701018A
Financial institution risk assessment system based on big data analysis
CN118761838A