A data transmission method and device, electronic equipment, storage medium and product

By using a three-stage verification mechanism for the unified access port and scheduling node of the tunnel proxy, the problem of unstable proxy service quality is solved, the success rate and efficiency of data transmission are improved, and the stability and security of data transmission are ensured.

CN122372236APending Publication Date: 2026-07-10CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD
Filing Date
2026-03-05
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

The quality of existing proxy services varies, resulting in a high failure rate for request sending, an inability to effectively manage proxy status, and an impact on the stability and efficiency of data transmission.

Method used

Client requests are received through a unified access port via a tunnel proxy. The scheduling node selects the target tunnel proxy from the proxy pool based on the main domain name and protocol. A three-stage verification method is used to manage the proxy status, including initial verification, in-use verification, and periodic re-verification, to ensure the stability of proxy quality and select the optimal tunnel proxy for data transmission.

Benefits of technology

It improves the success rate and efficiency of data transmission, avoids the use of unavailable or poor-quality proxies, ensures the accuracy and stability of data transmission, hides the client's real IP address, and enhances anonymity and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122372236A_ABST
    Figure CN122372236A_ABST
Patent Text Reader

Abstract

This application provides a data transmission method, data transmission device, electronic device, computer-readable storage medium, and computer program product. The method includes: receiving a request sent by a client through a unified access port of a tunnel proxy; wherein the request includes a main domain name and a protocol; selecting a target tunnel proxy from a proxy pool based on the main domain name and protocol through a scheduling node; wherein the status of each tunnel proxy in the proxy pool is managed by a service node through a three-stage verification method based on initial verification, in-use verification, and periodic re-verification; the target tunnel proxy is used to transmit data with a target server; sending a request to the target server through the target tunnel proxy; receiving request data fed back by the target server through the target tunnel proxy, and sending request data to the client.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to data transmission technology, and more particularly to a data transmission method, data transmission device, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] A proxy service is a common network service that acts as an intermediary between a client and a target server. The client sends a request through the proxy server, which then forwards the request to the target server. After the target server processes the request, the proxy server returns a response to the client.

[0003] Typically, proxy services are purchased from third parties and put into use directly. However, the quality of procured proxies varies greatly depending on the supplier, leading to a high request failure rate. Summary of the Invention

[0004] This application provides a data transmission method, a data transmission device, an electronic device, a computer-readable storage medium, and a computer program product.

[0005] The technical solution of this application embodiment is implemented as follows: This application provides a data transmission method, the method comprising: The request is received from the client through a unified access port via a tunnel proxy; the request includes the main domain name and protocol. The scheduling node selects a target tunnel proxy from the proxy pool based on the main domain name and the protocol; wherein, the status of each tunnel proxy in the proxy pool is managed by the service node through a three-stage verification method based on initial inspection, in-use verification, and periodic re-inspection; the target tunnel proxy is used to transmit data with the target server; The request is sent to the target server through the target tunnel proxy; The target tunnel proxy receives the request data from the target server and sends the request data to the client.

[0006] This application provides a data transmission device, including: The tunnel proxy uses a unified access port to receive requests sent by clients; these requests include the main domain name and the protocol. A scheduling node is used to select a target tunnel proxy from the proxy pool based on the main domain name and the protocol; wherein each tunnel proxy in the proxy pool is managed by the service node through a three-stage verification method based on initial verification, in-use verification, and periodic re-verification; the target tunnel proxy is used to transmit data with the target server; The target tunnel proxy is used to send the request to the target server; receive the request data fed back by the target server; and send the request data to the client.

[0007] This application provides an electronic device, including: Memory is used to store executable instructions or computer programs. The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the method provided in the embodiments of this application.

[0008] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the data transmission provided in this application when executed by a processor.

[0009] This application provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implements the data transmission provided in this application.

[0010] The embodiments of this application have the following beneficial effects: Requests sent by clients are received through a unified access port of the tunnel proxy; wherein the request includes the main domain name and protocol; a target tunnel proxy is selected from the proxy pool based on the main domain name and protocol by the scheduling node; wherein the status of each tunnel proxy in the proxy pool is managed by the service node through a three-stage verification method based on initial inspection, in-use verification, and periodic re-inspection; the target tunnel proxy is used to transmit data with the target server; a request is sent to the target server through the target tunnel proxy; the target tunnel proxy receives the request data fed back by the target server and sends the request data to the client; thus, this application can manage the proxy status through a three-stage verification method, ensuring the stability of the proxy quality in the proxy pool, and by selecting the optimal tunnel proxy from the proxy pool based on the main domain name and protocol by the scheduling node, improving the success rate and efficiency of requests, and avoiding the use of unavailable or poor-quality proxies. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating the data transmission method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the overall architecture of the network service of the tunnel proxy provided in the embodiments of this application; Figure 3 This is a schematic diagram of the load balancer selector of the scheduling node selecting an agent according to an embodiment of this application; Figure 4 This is a schematic diagram of the agent lifecycle management process provided in an embodiment of this application; Figure 5This is a schematic diagram of the proxy request lifecycle management process provided in the embodiments of this application; Figure 6 This is a schematic diagram of the structure of the data transmission device provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0013] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0014] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0015] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0016] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0017] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0018] Figure 1This is a flowchart illustrating the data transmission method provided in an embodiment of this application. The following will be combined with... Figure 1 The steps shown will be explained. It should be noted that... Figure 1 The method described uses an electronic device as the execution subject as an example. Figure 1 As shown, the method includes steps 101 to 104. Step 101: Receive requests sent by clients through the unified access port of the tunnel proxy; the requests include the main domain name and protocol.

[0019] In this embodiment, a fixed tunnel gateway (such as an Application Programming Interface (API) or a proxy server domain name) can be set as a unified access port. Clients initiate requests through this preset tunnel gateway, carrying information such as the main domain name and protocol in the request, without directly calling specific Internet Protocol Addresses (IP addresses). This entry point acts as an "intermediate layer," decoupling the user from the backend IP pool, so the client does not need to concern itself with the specific IP information of the backend. This achieves decoupling between the user and the backend IP pool, simplifying client operations; the client does not need to manually manage IPs, but only needs to initiate requests through a fixed entry point.

[0020] In practical applications, network programming techniques can be used to listen on this unified access port. When a client request arrives, the request data is received, and information such as the main domain name and protocol is extracted. For example, in Python, the socket library can be used to create a listening socket, listen on a specified port, and receive client requests. The unified access port of the tunnel proxy lays the foundation for subsequent dynamic IP scheduling and request processing, and the unified entry point facilitates centralized management and processing of requests.

[0021] Step 102: The scheduling node selects the target tunnel proxy from the proxy pool based on the main domain name and protocol; the status of each tunnel proxy in the proxy pool is managed by the service node through a three-stage verification method based on initial inspection verification, in-use verification and periodic re-inspection; the target tunnel proxy is used to transmit data with the target server.

[0022] In this embodiment, after receiving a request containing the main domain name and protocol, the scheduling node selects a target tunnel proxy from the proxy pool in real time. The selection criteria include, but are not limited to, one or more of IP availability, geographical location, latency, and bandwidth.

[0023] IP availability refers to avoiding IP address ranges that have been identified and blocked by the system when selecting proxy IPs. This application ensures IP availability through a three-stage verification method. The three-stage verification method includes: initial verification, in-use verification, and periodic re-verification. The initial verification is performed when the proxy is added to the proxy pool to check whether the IP is available, etc.; the in-use verification monitors in real time during the use of the proxy and determines whether the proxy is invalid if request failures are found; the periodic re-verification checks the proxy regularly to ensure that it is in an available state.

[0024] Among them, geolocation is used to match regional restrictions if the business has specific regional access requirements.

[0025] Among them, latency and bandwidth are used to optimize the routing algorithm and select proxies with low latency and high bandwidth.

[0026] In some embodiments, algorithms and data structures can be used to implement the agent selection logic. For example, a list of agents can be maintained, which records relevant information for each agent (such as availability, geographical location, latency, etc.). Agents can be sorted or filtered according to selection criteria to select the optimal target tunnel agent.

[0027] The method for selecting a target tunnel proxy based on this application can choose the optimal tunnel proxy from the proxy pool, improving the success rate and efficiency of requests and avoiding the use of unavailable or poor-quality proxies. Managing proxy status through a three-phase verification method ensures the stability of proxy quality in the proxy pool and reduces request failures caused by proxy issues.

[0028] Step 103: Send a request to the target server through the target tunnel proxy.

[0029] In this embodiment, after receiving the task assigned by the scheduling node, the target tunnel agent sends a request to the target server based on information such as the main domain name and protocol in the request.

[0030] In some embodiments, before sending a request, the request may be processed in the following ways, such as encapsulating request headers and setting request parameters according to protocol requirements.

[0031] This application initiates a request to the target server based on the main domain name and protocol information in the request, realizing the function of sending requests to the target server through a proxy, hiding the client's real IP address, and improving the client's anonymity and security. It can accurately send requests to the target server according to the protocol requirements, ensuring that the request is correctly received and processed by the target server.

[0032] Step 104: Receive the request data from the target server through the target tunnel proxy and send the request data to the client.

[0033] In this embodiment, the target tunnel proxy listens to the response of the target server. When it receives the request data from the target server, it sends the request data back to the client.

[0034] In some embodiments, after receiving the request data from the target server, the data can be processed (e.g., decrypted, if the data was encrypted during transmission). Then, the processed data is sent to the client through the established connection.

[0035] In this embodiment, relevant techniques in network programming can be used to send data, such as using the send method of the socket library in Python, or using pre-encapsulated network communication objects (such as the writer object in Asyncio) to send data.

[0036] This application receives request data from the target server via a target tunnel proxy and sends request data to the client, thus realizing the transmission of data from the target server to the client and enabling the client to obtain the target server's response data. The client's real IP address is hidden throughout the process, ensuring client privacy and security, while also guaranteeing the accuracy and stability of data transmission.

[0037] The data transmission method provided in this application receives requests sent by clients through a unified access port of a tunnel proxy. The request includes the main domain name and protocol. A scheduling node selects a target tunnel proxy from the proxy pool based on the main domain name and protocol. The status of each tunnel proxy in the proxy pool is managed by a service node using a three-stage verification method: initial verification, in-use verification, and periodic re-verification. The target tunnel proxy is used for data transmission with the target server. Requests are sent to the target server through the target tunnel proxy. The target tunnel proxy receives request data from the target server and sends request data back to the client. Therefore, this application can manage the proxy status through a three-stage verification method, ensuring the stability of the proxy quality in the proxy pool. Furthermore, by selecting the optimal tunnel proxy from the proxy pool based on the main domain name and protocol through the scheduling node, the success rate and efficiency of requests are improved, and the use of unavailable or poor-quality proxies is avoided.

[0038] In a feasible scenario, Figure 2 This is a schematic diagram of the overall architecture of the tunnel proxy network service provided in this application embodiment. This architecture, based on Python coroutines, implements a similar tunnel proxy network service, solving problems such as proxy quality, process visibility, and traffic controllability in one go. Figure 2As shown, this service uses a client / server (C / S) architecture, presented to the client as a tunnel proxy. The client only needs to connect to the tunnel proxy like a regular proxy to send Hypertext Transfer Protocol (HTTP) requests and receive responses. The entire request process is scheduled and audited by the server. This architecture includes: User layer: Multiple users connect to the tunnel proxy through "coroutine + port multiplexing" to achieve single-port multiple connection multiplexing.

[0039] Service layer: Includes service nodes (containing multiple service processes that communicate via "coroutines + sockets") and tunnel proxy. Service nodes are responsible for core logic such as request scheduling and IP switching, while tunnel proxy serves as a unified entry point for handling user connections and request forwarding.

[0040] Scheduling and monitoring layer: The scheduling node communicates with the service node through "coroutine + socket". The real-time statistics module collects data such as domain name, query per second (QPS), bandwidth, success rate and writes it to the log system.

[0041] In this embodiment, high concurrency capabilities can be achieved at a very low cost through the underlying technology of multi-process + coroutines. Real-world testing in a production environment shows that only 4 CPUs + 4GB of memory are needed to stably serve 1000+ QPS and 1000+ Mbps of bandwidth. In single-machine deployment, inter-process communication uses coroutines + Unix Domain Socket (UDS) to ensure performance as much as possible. UDS is an inter-process communication mechanism used only within the same host. If a multi-machine deployment is adopted, communication between different machines (nodes) switches to a host + port method using ordinary sockets. Both methods are abstracted by Asyncio to provide the same read / write object interface, allowing for seamless switching.

[0042] In some embodiments, step 102, where the scheduling node selects the target tunnel proxy from the proxy pool based on the primary domain name and protocol, can be implemented through the following steps: The scheduling node selects the target tunnel proxy by selecting available proxies from the proxy pool based on the main domain name and protocol, according to the proxy status and a strategy of iteratively selecting available proxies starting from a random offset position.

[0043] In this embodiment, the scheduling node first filters out a basic set of candidate proxies from the proxy pool that support the protocol and have a high success rate for requests to that main domain name, based on the request's main domain name and protocol. Within the index range of the candidate set, a starting offset is randomly generated. For example, if there are 10 proxies in the candidate set, a number between 0 and 9 is randomly generated as the starting position. Starting from the position corresponding to the random offset, the proxies in the candidate set are traversed sequentially, prioritizing proxies with a status of "available" and remaining available attempts not reaching the threshold. If no available proxy is found after one traversal, the emergency replenishment mechanism of the proxy pool is triggered, and a new compatible proxy is collected before re-executing the selection process. This application adopts a strategy of random offset + cyclic iteration, which can evenly distribute requests for the same main domain name to different proxies, avoiding the blocking of a single proxy by the target website due to excessive request frequency. At the same time, it can make full use of the available resources in the proxy pool and reduce retry costs. In addition, combined with real-time updated proxy status information, high-quality available proxies are prioritized, and a replenishment mechanism is triggered when no available proxy is available, ensuring that the proxy pool can always provide available resources, improving the overall success rate and stability of requests.

[0044] In some embodiments, after a target agent is selected, the remaining available times, the most recent usage time, and other status information of the agent are immediately updated and synchronized to the agent pool's storage system to provide the latest data for the next selection.

[0045] In some embodiments, before selecting a target tunnel proxy from the proxy pool, pre-processing operations can be performed: First, the proxies in the proxy pool are marked for protocol and domain name compatibility: proxies are pre-categorized according to their supported protocols (HTTP / Hypertext Transfer Protocol Secure, HTTPS), and the compatibility success rate of each proxy with common main domains is recorded and stored in an in-memory database such as Redis for easy and quick retrieval. Proxy status information is maintained: the liveness status, request success rate, response latency, and remaining available attempts of each proxy are updated in real time, serving as the core basis for proxy selection. Pre-matching of main domains and protocols prevents proxies with incompatible protocols or poor domain name compatibility from being invoked, reducing request error rates at the source; for example, HTTPS requests will not match proxies that only support HTTP.

[0046] Figure 3 This is a schematic diagram of the load balancer selecting an agent for the scheduling node provided in an embodiment of this application. Figure 3This diagram illustrates two proxy lists, Proxy List 1 and Proxy List 2, and the offsets of different domains within these lists. A domain name is a name used to identify and locate resources such as websites on the network. The offset is the starting position in the proxy list where proxy polling begins when a domain first selects a proxy; setting random offsets helps avoid congestion. Here, the offset position is the specific position number in the proxy list where the domain begins proxy polling. For example, if domain A's offset = 1, then proxy polling in Proxy List 1 will begin from position 1.

[0047] like Figure 3 As shown, load balancing during proxy selection is based on the main domain name. The aim is to maximize the interval between requests from each proxy to any given site, avoiding high-frequency access that could trigger risk control measures within the same timeframe. New proxies will be inserted into available slots or at the end of the queue, initially in the "under verification" state. like Figure 3 As shown, proxies 1-7 (representing proxy 7 in proxy list 1) and 2-1 (representing proxy 1 in proxy list 2) will change their status to HTTPS / HTTP / Eliminated after the initial verification. Eliminated proxies 1-2 (representing proxy 2 in proxy list 1), 1-3 (representing proxy 3 in proxy list 1), and 2-9 (representing proxy 9 in proxy list 2) are skipped during selection. During recycling, the reason for elimination is recorded and the proxies are set to idle. Idle nodes at the end of the list will be periodically pruned. Proxy 2-4 (representing proxy 4 in proxy list 2) is skipped after its penalty count is decremented by 1 during selection. When the penalty count reaches zero, the penalty ends and the proxy becomes available.

[0048] For the initial domain selection, a proxy list is first chosen, and then proxies are polled in a round-robin fashion based on a random offset to avoid congestion. In most cases, there is only one proxy in the list. The proxy selection logic is as follows: starting from the offset position, iterates in a loop. If a proxy in the penalty period is encountered, the penalty count is reduced by 1 until a usable proxy is found. Domain information is stored only at the offset to save storage costs, and the cache of domains that have not been used for a long time is cleaned up regularly.

[0049] The main domain name is used instead of the complete domain name in this application for two reasons. First, it is based on site considerations. Although the same site may have many hostnames / subdomains pointing to different IP addresses, it is very likely that the same risk control gateway device (such as a CC attack firewall) is used to count request frequency. Second, some sites have many subdomains, which will make the load balancing logic ineffective.

[0050] In some embodiments, the above data transmission method further includes the following steps: When each tunnel proxy first connects to the tunnel proxy's network service, the service node performs initial checks on HTTP availability and HTTPS availability respectively. Add tunnel proxies that pass the initial verification to the queue of the proxy pool.

[0051] In this embodiment, when a new tunnel proxy initiates an access request to the service node for the first time, the service node automatically triggers the verification process without manual intervention. The access request must carry basic proxy information, including IP address, port number, supported protocol types, etc., as initial parameters for verification.

[0052] Here, the HTTP availability check is performed as follows: The service node sends a standard HTTP request to the proxy (such as accessing a publicly accessible HTTP test domain like http: / / www.test.com), sets a timeout threshold (such as 10 seconds), and if a normal response with a 200 status code is received within the threshold and the response content meets expectations, the HTTP availability check is considered to have passed; if there is no response after the timeout, an error status code is returned, or the response content is abnormal, it is considered to have failed.

[0053] Here, HTTPS availability verification works as follows: The service node sends a standard HTTPS request to the proxy (such as accessing a publicly accessible HTTPS test domain like https: / / www.test.com), with a 10-second timeout threshold. If a normal response with a 200 status code is received and the content is normal, the HTTPS availability verification is considered successful; otherwise, it is considered unsuccessful. For proxies that only support HTTP, their HTTPS verification can be marked as "not supported," without affecting subsequent HTTP usage.

[0054] Among them, the verification result processing and proxy pool queuing are as follows: If a proxy passes both HTTP and HTTPS verification, it is marked as "full protocol support" and added to the full protocol availability queue of the proxy pool.

[0055] If only HTTP validation is passed, mark it as "HTTP Only" and add it to the HTTP-only queue of the proxy pool.

[0056] If only HTTPS verification is passed, mark it as "HTTPS only" and add it to the HTTPS-only queue of the proxy pool.

[0057] If both protocol verifications fail, access will be rejected directly, and the proxy information will be recorded in the abnormal proxy database, and no further access attempts will be made.

[0058] As described above, this application uses a dual-protocol availability check before access to directly filter out low-quality proxies that cannot provide services normally or are incompatible with the protocols. This avoids problems such as request failures and increased retry costs caused by these proxies entering the proxy pool, fundamentally improving the overall quality of the proxy pool. By categorizing proxies into queues based on their protocol support types, subsequent scheduling nodes can accurately match proxies in the corresponding queues based on the request's main domain name and protocol type, avoiding request errors caused by protocol mismatches and improving request success rates. Early elimination of unavailable proxies reduces the proportion of invalid resources in the proxy pool, lowers the pressure on scheduling nodes to select proxies, and avoids request interruptions due to proxy failures, improving the stability and reliability of the entire data transmission system. Categorizing and managing proxies with different protocol support allows each type of proxy to function effectively in its corresponding scenario, avoiding resource waste. For example, proxies that only support HTTP can be used only for HTTP request scenarios, maximizing the utilization of proxy resources.

[0059] In this embodiment, when each tunnel proxy first accesses the tunnel proxy's network service, the service node determines the new proxy's first access by any one of the following: (1) Active registration: The new tunnel agent submits a registration request to the service node through a preset interface or protocol, including its own network information (such as IP, port, supported protocols, etc.), triggering the service node's access process; (2) Service nodes identify and mark new proxy nodes by scanning a specified network range, receiving external data source pushes, or monitoring connection requests on specific ports; (3) When a service node detects a new proxy access request, it will assign a unique identifier to it and initialize metadata (such as source, generation time, etc.), and then start the initial inspection process to verify availability and complete the state switch from "new" to "available".

[0060] In some embodiments, the above data transmission method further includes the following steps: When any tunnel proxy in the proxy pool is selected as the target tunnel proxy, during the data transmission between the selected tunnel proxy and the target server, the service node obtains the request information output log corresponding to the request; the request information output log includes error information that occurs when each tunnel proxy is selected as the target tunnel proxy to process the request. The scheduling weights of the selected tunnel agents are adjusted by the service nodes based on error information.

[0061] In this embodiment, once the proxy is selected by the scheduling node and establishes a data transmission connection with the target server, the service node automatically starts log collection and monitoring. Throughout the entire request processing flow, the following key information is captured and recorded in real time: Proxy basic information: IP address, port number, protocol type, and queue; Request-related information: target domain name, request protocol, request time, and request path; Error information: specific error types and details such as request timeout, abnormal status codes (e.g., 4xx / 5xx), protocol incompatibility, connection interruption, and response content verification failure; Transmission performance data: response latency, request success rate, and data transmission volume.

[0062] In this embodiment, the collected error information is standardized, classified, and parsed by the service node: Errors are categorized by type: protocol-related (e.g., HTTPS requests using HTTP proxies), connection-related (e.g., connection timeouts), website blocking-related (e.g., 403 Forbidden), and performance-related (e.g., latency exceeding threshold). Key features are extracted: features that can pinpoint the problem are extracted from the error details, such as the target website domain, the time period in which the error occurred, and the frequency of the error.

[0063] Furthermore, based on the parsed error information, the scheduling weight of the agent is dynamically updated through preset weight adjustment rules.

[0064] Here are some scenarios for weight reduction: If a proxy experiences serious errors such as protocol incompatibility or connection interruption, its weight value is directly reduced, for example, from an initial weight of 10 to 5, to decrease the probability of it being scheduled. If a proxy experiences multiple request timeouts or response delays exceeding thresholds within a short period (e.g., within 1 hour), its weight is gradually reduced until it is removed from the high-frequency scheduling queue. If a proxy is blocked by the target website, causing request failures, its weight is directly reduced to the lowest level, and only gradually restored after a long period without error records.

[0065] Here, the weight recovery and enhancement scenario is as follows: If the proxy makes multiple successful requests consecutively and its performance metrics (latency, success rate) are excellent, its weight is gradually increased, thus increasing its scheduling priority. If the proxy experiences a few errors due to temporary network fluctuations, its weight automatically returns to its initial value after subsequent requests resume normal operation.

[0066] The adjusted weight information is synchronized in real time to the state storage system of the proxy pool (such as Redis) to ensure that the scheduling node can obtain the latest weight data.

[0067] In this embodiment, a trigger threshold for weight adjustment is set: for example, a single serious error directly triggers a weight reduction, while minor errors need to accumulate to a certain number before adjustment is triggered. A weight adjustment cycle is also set: for example, error information from the agent is summarized and analyzed every 10 minutes, and weights are adjusted in batches to avoid frequent weight fluctuations due to a single error.

[0068] As described above, this application, through error-information-based weight adjustment, prioritizes high-quality, low-error-rate proxies, while gradually eliminating or downgrading low-quality, high-error-rate proxies. This fundamentally reduces request failure rates and retry costs, improving overall data transmission success rates. Real-time weight adjustment based on proxies' performance allows for rapid response to fluctuations in proxies' quality. For example, if a proxy suddenly experiences a large number of errors due to network issues, its weight can be promptly reduced to prevent impacting overall request stability. Once the proxy recovers, its weight is automatically restored, fully utilizing available resources. Continuous error feedback and weight adjustment create a dynamic "survival of the fittest" management mechanism, continuously optimizing the resource quality of the proxy pool, reducing request interruptions and data transmission failures caused by proxy issues, and improving the reliability of the entire data transmission system. Automated error information collection and weight adjustment require no manual intervention, significantly reducing the operational costs of the proxy pool while ensuring the timeliness and accuracy of weight adjustments, avoiding the lag and subjectivity of manual judgment. By using the protocol type flag in the error message, we can accurately identify proxies with incompatible protocols, adjust their weights accordingly, and prevent such proxies from being used in mismatched request scenarios, thus solving the problem of request errors caused by protocol incompatibility at its root.

[0069] In some embodiments, the above data transmission method further includes the following steps: When the adjusted scheduling weight is less than the target threshold, a candidate tunnel proxy is selected from the proxy pool through the service node; the candidate tunnel proxy is used to transmit data with the target server, and the candidate tunnel proxy is different from the selected tunnel proxy. The selected tunnel agent is removed from the agent pool by the service node.

[0070] In this embodiment, when the scheduling weight of a proxy falls below a preset target threshold (e.g., weight value below 2), the service node automatically triggers a proxy replacement process: first, the low-weight proxy is locked, marked as "pending replacement," and its participation in new request scheduling is suspended; based on the protocol types supported by the original proxy (HTTP / HTTPS / all protocols) and the range of compatible main domains, a set of candidate proxies that meet the requirements is selected from the proxy pool to ensure that the candidate proxies can cover the service scenarios of the original proxy. This application continuously eliminates proxies with poor long-term performance and introduces high-quality candidate proxies through the automatic replacement of low-weight proxies, forming a closed-loop management of "survival of the fittest," and continuously improving the overall availability and stability of the proxy pool.

[0071] Here, the candidate proxy selection logic is as follows: Priority matching: Prioritize selection from the protocol-specific queue (e.g., the HTTP-only queue) to ensure protocol compatibility; if no available proxy is available in the queue, then filter from the full protocol support queue. Status verification: Candidate proxies must meet core conditions such as "available status," request success rate higher than 80%, and response latency lower than the threshold, excluding proxies in a blocked or faulty state. Load balancing supplementation: If there are many qualified proxies in the same queue, a random offset + iterative strategy (consistent with the proxy selection strategy in step 102) is used to select candidate proxies, avoiding excessive calls to a single proxy.

[0072] Here, proxy replacement and state synchronization work as follows: After selecting a candidate proxy, the service node removes the original low-weight proxy from the active queue of the proxy pool and synchronously updates the state storage system (such as Redis) of the proxy pool, marking the proxy as "obsolete". The candidate proxy is then marked as "active", and its basic information, weight value, and other data are synchronized to the proxy pool to ensure that the scheduling node can immediately obtain the latest proxy resources. If the original proxy still has unfinished requests, the service node will smoothly migrate these requests to the candidate proxy to avoid request interruption.

[0073] In this embodiment, an exception fallback mechanism can also be provided: if no suitable candidate proxy can be found in the proxy pool, a proxy supplementation and collection process is triggered. A new adapted proxy is obtained through a web crawler or a third-party proxy collection interface, verified, and added to the proxy pool before the replacement operation is performed. If the original proxy's weight returns to normal (above the threshold) within a short period of time after being marked as "to be replaced," the replacement process is canceled, and its active state is restored.

[0074] As described above, this application employs designs such as smooth migration of incomplete requests and an exception fallback mechanism to prevent request interruption or failure due to proxy replacement, ensuring the continuity and reliability of data transmission. The selection of candidate proxies strictly matches the protocol and domain name compatibility range of the original proxy, avoiding issues such as protocol incompatibility and poor domain name compatibility caused by proxy replacement, thus maintaining the success rate of requests. Invalid proxies are promptly eliminated, ensuring the proxy pool remains in a highly efficient and available state, maximizing the utilization of high-quality proxy resources, and reducing retry costs and resource waste caused by proxy quality issues.

[0075] In some embodiments, the above data transmission method further includes the following steps: The service node performs periodic re-inspections on each tunnel proxy in the proxy pool; the periodic re-inspection consists of initial verification and in-use verification performed at preset time intervals.

[0076] In this embodiment, the service node automatically triggers a scheduled re-inspection task for the entire proxy pool at preset time intervals (e.g., every 24 hours). Different re-inspection cycles can also be set for proxies with different protocols and quality levels: for example, high-quality proxies supporting all protocols are re-inspected every 48 hours, proxies supporting only a single protocol are re-inspected every 24 hours, and low-weight proxies are re-inspected every 12 hours. The re-inspection tasks are executed in batches and in parallel to avoid excessive system resource consumption for individual proxy re-inspections. This application, through periodic re-inspections, promptly detects problems such as IP failures, protocol compatibility changes, and website blocking that occur during proxy use, preventing low-quality proxies from remaining in the proxy pool for extended periods and maintaining the availability of the proxy pool dynamically.

[0077] In this embodiment, a two-dimensional verification process is implemented through service nodes: Initial verification reuse: The HTTP / HTTPS availability verification logic from the initial access is reused. A standard HTTP / HTTPS request is sent to the proxy, a 10-second timeout threshold is set, and the basic availability of the proxy is judged by the response status code and content integrity. If the verification fails, it is directly marked as "unavailable". Verification during use is newly added: For proxies already running in the pool, scenario-based verification is added: Accessing target websites in different regions and industries verifies the proxy's cross-scenario availability; simulating high-concurrency requests (such as initiating 100 requests simultaneously) to test the proxy's load-bearing capacity; verifying the authenticity of the proxy's IP address to avoid IP spoofing or blocking.

[0078] In this embodiment, the service node can also perform tiered processing of the verification results: if a proxy passes both the initial check and the in-use check, it is marked as "high-quality available," maintaining or increasing its scheduling weight; if it only passes the initial check, it is marked as "basic available," maintaining its original scheduling weight but increasing the re-check frequency; if the initial check fails, it is directly removed from the proxy pool and synchronously recorded in the abnormal proxy database; if the in-use check fails but the initial check passes, it is marked as "restricted available," reducing its scheduling weight and scheduling it only in low-concurrency scenarios, while shortening the re-check cycle, and restoring its weight after passing three consecutive re-checks. This application's tiered processing based on re-check results allows high-quality proxies to obtain higher scheduling priority, restricts available proxies to be used only in appropriate scenarios, further improving request success rate and transmission efficiency.

[0079] In this embodiment, the results of each re-inspection (including verification time, verification items, pass status, and error details) are stored in the agent status database (such as MySQL) and synchronized in real time to the caching system of the scheduling node (such as Redis) to ensure that the scheduling node can obtain the latest agent status information and provide data support for subsequent agent selection.

[0080] Figure 4 This is a flowchart illustrating the agent lifecycle management process provided in an embodiment of this application, as shown below. Figure 4 As shown, when a new proxy enters the server, it first undergoes two rounds of availability checks, i.e., two initial checks (HTTP / HTTPS), to filter HTTPS and HTTP availability and mark the available protocols. To save bandwidth and improve performance, the validation interface uses the generate204 interface with no content.

[0081] After successful verification, the proxy supplements its metadata information (including but not limited to one or more of the following: source, expiration time, status, number of requests, number of successes, number of penalties, verification time, and generation time), and adds it to the proxy pool queue for allocation and use. The proxy metadata state machine switching list is as follows: Add, Initial Check, Re-check, Penalty, Available, Expired.

[0082] The load balancer selects the next proxy to be used for the specified domain based on the input parameters (main domain name + protocol), sends requests and receives responses, records request information and outputs logs. For errors thrown when requesting a proxy (such as no downlink traffic, connection timeout, connection rejection, authentication failure, or other reasons), it performs graded demotion and demotion. If the elimination threshold is reached, the proxy is directly eliminated.

[0083] Therefore, each agent has three verification methods in its lifecycle: first, an initial check to determine whether it can enter the agent pool queue; second, a real-time and dynamic adjustment of scheduling weights based on connectivity and response during use; and third, a periodic re-check to further improve the overall quality.

[0084] This application, through a three-tiered proxy availability verification standard and a dynamically adjusted weight audit strategy, combined with the feature marking function in use, initially solves the problems of timeliness and cost waste in verifying proxy availability. In large-scale data collection task scenarios, it efficiently verifies a large number of proxies and automatically downgrades or eliminates unstable proxies, thereby improving the overall quality and efficiency of proxy usage.

[0085] Figure 5 This is a schematic diagram of the management process for the proxy request lifecycle provided in the embodiments of this application, such as... Figure 5 As shown, the process includes the following: First, the client establishes a connection with the tunnel proxy and sends a request. The server reads the first line of the request header and extracts the main domain name and protocol.

[0086] The input parameters of the flow control module of the scheduling node include the original request information such as the domain name and the client address.

[0087] Secondly, the traffic control module performs fine-grained control of traffic based on custom strategies.

[0088] The circuit breaker in the flow control module is responsible for determining whether the domain name blacklist and the rate limiter have an idle token. If the circuit breaker condition is triggered, the connection is disconnected and the reason is recorded in the log.

[0089] The flow control module's rate limiter obtains tokens from the token bucket, determines whether it needs to control the request interval, and if cooling is required, it waits in place (while checking if the connection is broken) until it can send a request and returns the token.

[0090] The traffic control module requests an available proxy for the domain from the scheduling service based on the main domain name and protocol. It may encounter problems such as no proxy available, which will not be elaborated here. After obtaining a proxy, it attempts to establish a connection. If the connection fails, it requests an available proxy again with the failure information until the connection is successful or the client has been disconnected.

[0091] In this embodiment, the reader and writer objects initialized using Python's built-in Asyncio framework are used to create upload and download pipeline objects with the existing reader and writer objects on the client side for data transmission. During pipeline initialization, request headers can also be further encapsulated using custom functions, such as modifying the headers to update cookie or authorization fields for sites requiring login.

[0092] According to the message characteristics of the HTTP 1.0 / 1.1 protocol, if the Pipeline switches from download to upload status during this connection, it indicates that the client has initiated a new request, and the request count is incremented by 1. Wait for the client to disconnect, record the original information of this connection, such as uplink traffic, downlink traffic, and request count, and return it to the scheduling service along with the proxy information.

[0093] Before the client disconnects, it can send a specific HTTP request to the proxy service during the current connection to mark the proxy IP used last time. For example, it can trigger anti-crawling risk control so that requests for this domain name will skip the proxy IP in the next scheduling.

[0094] The scheduling service collects information on the proxy usage, adjusts the proxy weight, and generates single request logs and minute-level statistical logs, including request counts, uplink traffic, downlink traffic, and domain information. As can be seen from the above, this application solves the usability problem of large-scale proxy requests through tunnel proxy. It addresses the stability problem of proxy quality through a three-layer verification method. By collecting statistical information for each connection, it solves the problem of unobservable proxy usage. Through a traffic controller and load balancing scheduling mechanism, it solves the problem of uncontrollable traffic. By sending HTTP requests within the same connection, it marks the previously used IP and domain name, skipping proxy addresses that have already triggered anti-scraping measures.

[0095] This application has the following beneficial effects: This application proposes a highly customizable tunnel proxy implementation method. By analyzing the traffic characteristics of HTTP requests, it can achieve functions such as load balancing, site traffic control, and customized proxy selection for IP requests to the same site. In large-scale data collection scenarios, it avoids the risk of uncontrolled request volume to the target website. Furthermore, it standardizes the statistical information of different entities, so that each proxy, each request, and each domain name has formatted statistical results within its lifecycle. This makes the originally "black box" tunnel proxy more observable, reduces debugging costs, and increases the dimensions of real-time monitoring indicators.

[0096] This application, through a three-tiered proxy availability verification standard and a dynamically adjusted weight audit strategy, combined with the feature marking function in use, initially solves the problems of timeliness and cost waste in verifying proxy availability. In large-scale data collection task scenarios, it efficiently verifies a large number of proxies and automatically downgrades or eliminates unstable proxies, thereby improving the overall quality and efficiency of proxy usage.

[0097] With the proxy service implemented in this application, the number and availability of proxies provided by different proxy providers can be monitored in real time, and the cause of the problem can be quickly located, such as insufficient bandwidth of a certain provider and lack of HTTPS support for some proxies. After the problem is resolved, the availability rate increases from 45% to 99%.

[0098] Using the proxy service implemented in this application, it was quickly located that the distributed crawler program had more than 1 million instantaneous visits to sites A and B at 0:00 every day, which far exceeded the safe request frequency. By setting the site's default maximum QPS, the excessively fast requests were automatically circuit-broken and alarm emails were sent. Further investigation revealed that the production accident was caused by infinite retries due to a code problem.

[0099] This application also provides a data transmission device for use in electronic devices, such as... Figure 6 As shown, the data transmission device includes: The tunnel proxy uses a unified access port of 601 to receive requests sent by clients; these requests include the main domain name and the protocol. Scheduling node 602 is used to select a target tunnel proxy from the proxy pool based on the main domain name and protocol. Each tunnel proxy in the proxy pool is managed by the service node through a three-stage verification method of initial verification, in-use verification, and periodic re-verification. The target tunnel proxy is used to transmit data with the target server. The target tunnel proxy 603 is used to send requests to the target server; receive request data from the target server and send request data back to the client.

[0100] In some embodiments, scheduling node 602 is used to select a target tunnel proxy from the proxy pool based on the primary domain name and protocol, according to the proxy status and following a strategy of iteratively selecting available proxies starting from a random offset position. In some embodiments, the data transmission apparatus further includes: Service node 604 is used to perform initial checks on HTTP availability and HTTPS availability when each tunnel proxy first accesses the network service of the tunnel proxy; and to add the tunnel proxies that pass the initial checks to the queue of the proxy pool.

[0101] In some embodiments, service node 604 is used to obtain request information output logs corresponding to the request during the data transmission process between the selected tunnel proxy and the target server when any tunnel proxy in the proxy pool is selected as the target tunnel proxy; wherein, the request information output logs include error information that occurs when each tunnel proxy is selected as the target tunnel proxy to process the request; Service node 604 is used to adjust the scheduling weight of the selected tunnel agent based on error information.

[0102] In some embodiments, service node 604 is used to select candidate tunnel proxies from the proxy pool when the adjusted scheduling weight is less than the target threshold; wherein, the candidate tunnel proxies are used to transmit data with the target server, and the candidate tunnel proxies are different from the selected tunnel proxies; Service node 604 is used to remove the selected tunnel agent from the agent pool.

[0103] In some embodiments, service node 604 is used to perform periodic re-inspections on each tunnel proxy in the proxy pool; wherein, the periodic re-inspection is to perform initial inspection and in-use inspection at preset time intervals.

[0104] This application also provides an electronic device, such as... Figure 7 As shown, electronic device 700 includes: Memory 701 is used to store computer-executable instructions or computer programs; When processor 702 executes computer-executable instructions or computer programs stored in memory 701, it performs the following steps: The request is received from the client through a unified access port via a tunnel proxy; the request includes the main domain name and protocol. The scheduling node selects the target tunnel proxy from the proxy pool based on the main domain name and protocol. The status of each tunnel proxy in the proxy pool is managed by the service node through a three-stage verification method: initial check, in-use check, and periodic re-check. The target tunnel proxy is used to transmit data with the target server. Send a request to the target server through the target tunnel proxy; The target tunnel proxy receives request data from the target server and sends request data to the client.

[0105] In some embodiments, processor 702 is configured to perform the following steps when selecting a target tunnel proxy from the proxy pool via a scheduling node based on the primary domain name and protocol: The scheduling node selects the target tunnel proxy by selecting available proxies from the proxy pool based on the main domain name and protocol, according to the proxy status and a strategy of iteratively selecting available proxies starting from a random offset position.

[0106] In some embodiments, when processor 702 executes computer-executable instructions or computer programs stored in memory 701, it performs the following steps: When each tunnel proxy first connects to the tunnel proxy's network service, the service node performs initial checks on HTTP availability and HTTPS availability respectively. The service node adds tunnel proxies that pass the initial verification to the queue of the proxy pool.

[0107] In some embodiments, when processor 702 executes computer-executable instructions or computer programs stored in memory 701, it performs the following steps: When any tunnel proxy in the proxy pool is selected as the target tunnel proxy, during the data transmission between the selected tunnel proxy and the target server, the service node obtains the request information output log corresponding to the request; the request information output log includes error information that occurs when each tunnel proxy is selected as the target tunnel proxy to process the request. The scheduling weights of the selected tunnel agents are adjusted by the service nodes based on error information.

[0108] In some embodiments, when processor 702 executes computer-executable instructions or computer programs stored in memory 701, it performs the following steps: When the adjusted scheduling weight is less than the target threshold, a candidate tunnel proxy is selected from the proxy pool through the service node; the candidate tunnel proxy is used to transmit data with the target server, and the candidate tunnel proxy is different from the selected tunnel proxy. The selected tunnel agent is removed from the agent pool by the service node.

[0109] In some embodiments, when processor 702 executes computer-executable instructions or computer programs stored in memory 701, it performs the following steps: The service node performs periodic re-inspections on each tunnel proxy in the proxy pool; the periodic re-inspection consists of initial verification and in-use verification performed at preset time intervals.

[0110] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the data transmission method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0111] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the data transmission method described above in this application.

[0112] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the data transmission method provided in this application. For example, ... Figure 1 The data transmission method is shown.

[0113] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0114] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0115] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0116] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0117] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A data transmission method, characterized in that, The method includes: The request is received from the client through a unified access port via a tunnel proxy; the request includes the main domain name and protocol. The scheduling node selects a target tunnel proxy from the proxy pool based on the main domain name and the protocol; wherein, the status of each tunnel proxy in the proxy pool is managed by the service node through a three-stage verification method based on initial inspection, in-use verification, and periodic re-inspection; the target tunnel proxy is used to transmit data with the target server; The request is sent to the target server through the target tunnel proxy; The target tunnel proxy receives the request data from the target server and sends the request data to the client.

2. The method according to claim 1, characterized in that, The step of selecting a target tunnel proxy from the proxy pool through a scheduling node based on the main domain name and the protocol includes: The scheduling node selects the target tunnel proxy from the proxy pool based on the main domain name and the protocol, according to the proxy status and following a strategy of iteratively selecting available proxies starting from a random offset position.

3. The method according to claim 1, characterized in that, The method further includes: When each tunnel agent first accesses the tunnel agent's network service, the service node performs initial checks on the availability of Hypertext Transfer Protocol (HTTP) and the availability of Hypertext Transfer Security Protocol (HTTP). The service node adds tunnel proxies that pass the initial verification to the queue of the proxy pool.

4. The method according to claim 3, characterized in that, The method further includes: When any tunnel proxy in the proxy pool is selected as the target tunnel proxy, during the data transmission between the selected tunnel proxy and the target server, the service node obtains the request information output log corresponding to the request; wherein, the request information output log includes error information that occurs when each tunnel proxy is selected as the target tunnel proxy to process the request; The service node adjusts the scheduling weight of the selected tunnel agent based on the error information.

5. The method according to claim 4, characterized in that, The method further includes: When the adjusted scheduling weight is less than the target threshold, a candidate tunnel proxy is selected from the proxy pool through the service node; wherein, the candidate tunnel proxy is used to transmit data with the target server, and the candidate tunnel proxy is different from the selected tunnel proxy; The selected tunnel agent is removed from the agent pool by the service node.

6. The method according to claim 5, characterized in that, The method further includes: The service node performs periodic re-inspections on each tunnel proxy in the proxy pool; wherein, the periodic re-inspection consists of performing the initial inspection and the in-use inspection at preset time intervals.

7. A data transmission device, characterized in that, The device includes: The tunnel proxy uses a unified access port to receive requests sent by clients; these requests include the main domain name and the protocol. A scheduling node is used to select a target tunnel proxy from the proxy pool based on the main domain name and the protocol; wherein each tunnel proxy in the proxy pool is managed by the service node through a three-stage verification method based on initial verification, in-use verification, and periodic re-verification; the target tunnel proxy is used to transmit data with the target server; The target tunnel proxy is used to send the request to the target server; receive the request data fed back by the target server; and send the request data to the client.

8. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.