A method and system for optimizing PCB test needle mark consistency based on reinforcement learning
Patent Information
- Application Number
- CN202611327299.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-31
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]针对上述存在的技术不足,本发明的目的是提出一种基于强化学习的PCB测试针痕一致性优化方法,旨在解决现有技术中采用固定权重的线性加权多目标奖励造成优化偏向,尤其是在PCB板材与探针状态持续变化条件下,固定权重无法动态平衡接触稳定度、针痕深度与失败风险权重配比,容易陷入对单一目标过度追求导致针痕一致性退化的技术问题
本发明通过采集接触阶段位移信号、实际针痕深度和电测试结果,利用香农熵量化接触稳定度、归一化深度偏移指数和连续化失败风险惩罚因子构建复合工况熵向量,全面表征探针-焊盘接触界面的瞬时状态,使工况辨识依据更敏感、更准确。在此基础上,采用双层自适应模糊推理机制生成动态权重向量,在发生严重异常时快速调整权重以保障测试安全,在正常工况波动时通过离线训练的神经模糊网络实现连续平滑的多目标权重过渡,从而使复合奖励信号始终与当前工况的优化优先级匹配,解决了固定权重奖励函数引起的优化偏置问题,能够引导策略网络在探针全生命周期内稳定控制针痕深度并维持高测试通过率。
Smart Images

Figure CN122819002A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data prediction and control technology, and in particular to a method and system for optimizing the consistency of PCB test pin marks based on reinforcement learning. Background Technology
[0002] Currently, in the electrical performance testing stage of mass PCB manufacturing, contact testing methods such as flying probe testing are widely used. The testing process requires ensuring stable contact between the probe and the pad while controlling the probe mark depth to 5-20μm. Within the target range, the aim is to reduce damage to the pads and maintain test consistency. To address the impact of probe state drift and board characteristic fluctuations on pinning parameters, the industry has introduced reinforcement learning for online closed-loop control of pinning pressure or displacement. Existing technologies typically linearly weight multiple optimization objectives such as contact stability, pin mark depth, and test failure rate with fixed weights to form a single reward function. The agent updates the policy network based on this reward after each pinning cycle, attempting to continuously optimize pinning behavior in batch testing.
[0003] For example, in actual mass testing, the PCB board type (such as the significant difference in overshoot sensitivity between ceramic and flexible substrates) and probe wear state will continuously change over time. When the probe enters the middle stage of wear, the energy release pattern at the contact interface becomes more chaotic, leading to a gradual increase in contact stability entropy and a prolonged contact impedance transient recovery time. If a preset fixed low contact stability weight is still used at this time, the reward function will guide the agent to continue to over-focus on the pin mark depth target, causing the probe to be tested before it has fully established stable contact, which can easily lead to an increase in false test rate or even hard failure. Conversely, if the pin mark depth weight is blindly increased, it will also drive the probe to press down excessively when the contact quality is still good, accelerating wear and causing the pin mark depth to exceed the tolerance. The fixed weight reward mechanism cannot recognize these dynamic changes in operating conditions, making the reinforcement learning strategy prone to over-pursuing a certain goal, thereby causing the overall pin mark consistency to degrade.
[0004] Therefore, there is an urgent need for a multi-objective dynamic reward mechanism that can automatically balance the relationship between contact stability, pin mark depth, and test failure risk in large-scale PCB testing scenarios where probe status and board characteristics are constantly changing. This mechanism would enable the reinforcement learning agent to adaptively adjust the weight ratio of each objective in the reward based on real-time feedback of contact stability fluctuations, pin mark depth offsets, and test failure risk, thereby overcoming the optimization bias problem caused by fixed-weight rewards and improving pin mark consistency and contact reliability in batch testing. Summary of the Invention
[0005] To address the aforementioned technical shortcomings, the present invention aims to propose a reinforcement learning-based optimization method for PCB test pin mark consistency. This method addresses the optimization bias caused by the use of linear weighted multi-objective rewards with fixed weights in existing technologies. In particular, under conditions of continuous changes in the PCB board material and probe state, fixed weights cannot dynamically balance the weight ratio of contact stability, pin mark depth, and failure risk, and are prone to falling into the technical problem of excessive pursuit of a single objective leading to pin mark consistency degradation.
[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: The present invention provides a PCB test pin mark consistency optimization method based on reinforcement learning.
[0007] The reinforcement learning-based PCB test pin mark consistency optimization method includes: Step S10: Obtain the multidimensional evaluation data of the previous acupuncture cycle, and perform the composite working condition characterization task based on the multidimensional evaluation data using entropy measurement and risk quantification mechanism, and output the composite working condition characterization vector. Step S20: Based on the composite working condition representation vector, a hierarchical adaptive fuzzy inference mechanism is used to perform the dynamic weight vector generation task and output the dynamic weight vector; Step S30: Based on the composite working condition representation vector and the dynamic weight vector, a weighted reward aggregation mechanism is used to perform the composite reward signal generation task, and the composite reward signal is output; Step S40: Based on the composite reward signal and the dynamic weight vector, a proportional attribution decomposition mechanism is used to perform a reward attribution calculation task, and a reward attribution vector is output. Step S50: Based on the reward attribution vector, the reinforcement learning strategy network update task is executed using the attribution-guided gradient mechanism, and the PCB test pin mark consistency optimization parameters are output.
[0008] Preferably, step S10, which involves acquiring multidimensional evaluation data from the previous acupuncture cycle, performing a composite working condition characterization task based on the multidimensional evaluation data using an entropy measurement and risk quantification mechanism, and outputting a composite working condition characterization vector, specifically includes: Step S101: Read the high-frequency displacement sensor data of the previous acupuncture cycle from the multidimensional evaluation data; perform Hilbert transform on the displacement signal corresponding to the high-frequency displacement sensor data within a preset time window to obtain an instantaneous amplitude sequence; extract peak points from the instantaneous amplitude sequence and divide the peak points into a preset number of amplitude intervals using the K-means clustering method; calculate the information entropy according to the proportion of the number of peak points in each amplitude interval to the total number of peak points to obtain the contact stability entropy, wherein the contact stability entropy satisfies:
[0009] in, This represents the contact stability entropy; This represents the amplitude range of the preset number; Indicates the amplitude interval index, and ; Indicates the first A range of amplitude values; Indicates falling into the first The number of peak points within each amplitude range; This indicates the total number of peak points; This represents the natural logarithm operation; when the total number of peak points is 0, the contact stability entropy is set to a preset default entropy value; Step S102: Read the actual needle depth of the previous needle-piercing cycle from the multidimensional evaluation data, obtain the preset material impact sensitivity correction factor corresponding to the current PCB material type, and calculate the normalized needle depth offset index based on the actual needle depth, the preset target needle mark depth, and the preset material impact sensitivity correction factor. The normalized needle depth offset index satisfies:
[0010] in, This represents the normalized needle mark depth offset index; This represents the preset material impact sensitivity correction factor; This indicates the actual needle insertion depth; This indicates the preset target pin mark depth, and the preset target pin mark depth is a value greater than 0; when the current PCB material type is a ceramic substrate, the preset material impact sensitivity correction factor is greater than 1; when the current PCB material type is a flexible substrate, the preset material impact sensitivity correction factor is less than 1. Step S103: Read the pass / fail flag and contact resistance transient recovery time of each test channel from the multidimensional evaluation data, and determine the failure risk factor based on the pass / fail flag and the contact resistance transient recovery time; wherein, when all the test channels pass and the contact resistance transient recovery time is less than a preset recovery time threshold, the failure risk factor is set to 0; when all the test channels pass and the contact resistance transient recovery time is not less than the preset recovery time threshold, the failure risk factor is set to 0.3; when at least one of the test channels fails, the failure risk factor is set to 1; the preset recovery time threshold is 10 ms; Step S104: According to a preset arrangement order, combine the contact stability entropy, the normalized needle mark depth offset index, and the failure risk factor into the composite working condition characterization vector, wherein the composite working condition characterization vector satisfies:
[0011] in, This represents the composite working condition characterization vector; Indicates the current moment of strategy decision-making; This refers to the failure risk factor.
[0012] Preferably, step S20, which involves using a hierarchical adaptive fuzzy inference mechanism to generate a dynamic weight vector based on the composite working condition representation vector and outputting the dynamic weight vector, specifically includes: Step S201: Input the failure risk factor and the normalized pin mark depth offset index into a preset coarse-grained fuzzy rule response layer to obtain a coarse-grained weight vector; wherein, the coarse-grained weight vector includes three coarse-grained weight components arranged in the order of contact stability, pin mark depth, and failure risk; the preset risk threshold is 0.7, and the preset depth offset threshold is 80%; when the failure risk factor is greater than the preset risk threshold, the coarse-grained weight vector is set to... When the failure risk factor is not greater than the preset risk threshold, and the normalized needle mark depth offset index is greater than the preset depth offset threshold, the coarse-grained weight vector is set to... ; Step S202: When the failure risk factor is not greater than the preset risk threshold and the normalized pin mark depth offset index is not greater than the preset depth offset threshold, the contact stability entropy, the normalized pin mark depth offset index, the failure risk factor and the contact resistance transient recovery time are input into the pre-trained adaptive neural fuzzy inference network to obtain a fine-grained weight vector. Step S203: Determine the fusion coefficient based on the failure risk factor, and use the fusion coefficient to perform weighted fusion of the coarse-grained weight vector and the fine-grained weight vector to obtain the dynamic weight vector, wherein the dynamic weight vector satisfies:
[0013] in, Represents the dynamic weight vector, and ; , and These represent the dynamic weight components corresponding to contact stability, pin mark depth, and failure risk, respectively. This represents the coarse-grained weight vector; This represents the fine-grained weight vector; This represents the fusion coefficient, and the fusion coefficient is equal to the failure risk factor.
[0014] Preferably, in step S202, the adaptive neural fuzzy inference network is obtained through offline training on historical best needle insertion data. The adaptive neural fuzzy inference network sets multiple Gaussian membership functions for each input variable and constructs a fuzzy rule base based on the historical best needle insertion data. The adaptive neural fuzzy inference network is used to learn the nonlinear continuous mapping relationship between the contact stability entropy, the contact resistance transient recovery time, and the fine-grained weight vector. When the contact stability entropy increases to 0.6 and the contact resistance transient recovery time increases to 15 ms, the adaptive neural fuzzy inference network continuously adjusts the fine-grained weight component corresponding to the contact stability from 0.3 to 0.5 and continuously adjusts the fine-grained weight component corresponding to the needle mark depth from 0.5 to 0.3.
[0015] Preferably, step S30, which involves using a weighted reward aggregation mechanism to generate a composite reward signal based on the composite working condition representation vector and the dynamic weight vector, and then outputting the composite reward signal, specifically includes: Step S301: Calculate the contact stability score based on the contact stability entropy and the preset entropy-stability scoring function to obtain the contact stability score. The contact stability score is negatively correlated with the contact stability entropy. Step S302: Calculate the needle mark depth score based on the normalized needle mark depth offset index and the preset needle mark depth scoring rule to obtain the needle mark depth score. Wherein, when the actual needle insertion depth is within the preset target depth range of 5μm to 20μm, and the deviation between the actual needle insertion depth and the preset target needle mark depth decreases, the needle mark depth score increases. Step S303: Calculate the switch-type failure risk score based on the failure risk factor and the pass / fail flag to obtain the failure risk score. When all the test channels pass, the failure risk score is set to a preset positive reward value; when at least one of the test channels fails, the failure risk score is set to a preset negative reward value. Step S304: Using the three dynamic weight components in the dynamic weight vector, the contact stability score, the pin mark depth score, and the failure risk score are weighted and summed to obtain the total reward value. The contact stability score, pin mark depth score, failure risk score, and total reward value are then encapsulated into the composite reward signal, wherein the total reward value satisfies:
[0016] in, This represents the total reward value.
[0017] Preferably, step S40, which involves performing a reward attribution calculation task based on the composite reward signal and the dynamic weight vector using a proportional attribution decomposition mechanism, and outputting a reward attribution vector, specifically includes: Step S401: Extract the contact stability score, the pin mark depth score, and the failure risk score from the composite reward signal, and multiply the contact stability score, the pin mark depth score, and the failure risk score by the corresponding dynamic weight components in the dynamic weight vector to obtain three weighted reward contribution values. Step S402: Perform signed normalization processing on the three weighted reward contribution values to obtain the contact stability attribution component, the pin mark depth attribution component, and the failure risk attribution component, wherein the attribution components satisfy:
[0018] in, Indicates the first Each attribution component; Indicates the scoring item number, and when When corresponding to the contact stability score, when When corresponding to the needle mark depth score, when The corresponding failure risk score; Indicates the first The rating value corresponding to each rating item; Indicates the first The rating value corresponding to each rating item; Indicates the summation sequence number; This represents the minimum positive number that is preset to avoid a denominator of 0. and These represent the i-th and j-th dynamic weight components, respectively. Step S403: Combine the contact stability attribution component, the pin mark depth attribution component, and the failure risk attribution component in the order of arrangement to obtain the reward attribution vector.
[0019] Preferably, step S50, which involves executing the reinforcement learning policy network update task based on the reward attribution vector using an attribution-guided gradient mechanism and outputting PCB test pin mark consistency optimization parameters, specifically includes: Step S501: Determine the contact stability attribution component, pinhole depth attribution component, and failure risk attribution component in the reward attribution vector as gradient-guided weights for the contact stability sub-objective, pinhole depth sub-objective, and failure risk sub-objective, respectively. Step S502: For each sub-objective, determine the sub-objective advantage estimate based on the deviation of the corresponding score value from the preset sub-objective baseline estimate, and calculate the gradient of the policy network objective function with respect to the policy network parameters according to the following formula:
[0020] in, Represents the objective function of the policy network; Indicates the policy network parameters; This represents the strategy state composed of the composite operating condition representation vector; This indicates the adjustment action of the control parameters output by the strategy network for the next injection cycle; This represents the conditional probability that the policy network will output the control parameter adjustment action under the policy state; Indicates the first Sub-objective advantage estimation for each sub-objective; This represents the gradient operation applied to the network parameters of the policy; This represents the expected value operation at the moment of strategy decision-making; Step S503: Along the direction that increases the objective function of the policy network, update the policy network parameters according to the gradient of the objective function of the policy network relative to the policy network parameters to obtain the PCB test pin mark consistency optimization parameters.
[0021] This invention also provides a reinforcement learning-based PCB test pin mark consistency optimization system, comprising: The composite working condition characterization module is used to acquire multi-dimensional evaluation data of the previous needle puncture cycle, and to perform the composite working condition characterization task based on the multi-dimensional evaluation data using an entropy measurement and risk quantification mechanism, and output the composite working condition characterization vector. The dynamic weight generation module is used to perform a dynamic weight vector generation task based on the composite working condition representation vector using a hierarchical adaptive fuzzy inference mechanism, and outputs a dynamic weight vector. The composite reward generation module is used to perform a composite reward signal generation task based on the composite working condition representation vector and the dynamic weight vector using a weighted reward aggregation mechanism, and output a composite reward signal. The reward attribution decomposition module is used to perform a reward attribution calculation task based on the composite reward signal and the dynamic weight vector using a proportional attribution decomposition mechanism, and output a reward attribution vector. The policy network update module is used to perform a reinforcement learning policy network update task based on the reward attribution vector using an attribution-guided policy gradient mechanism, and output PCB test pin mark consistency optimization parameters.
[0022] The present invention also provides a PCB test pin mark consistency optimization device based on reinforcement learning. The PCB test pin mark consistency optimization device based on reinforcement learning includes: a memory, a processor, and a PCB test pin mark consistency optimization program based on reinforcement learning stored in the memory and executable on the processor. When the PCB test pin mark consistency optimization program based on reinforcement learning is executed by the processor, it implements the above-described method.
[0023] The present invention also provides a computer program product, the computer program product including a reinforcement learning-based PCB test pin mark consistency optimization program, which implements the above method when executed by a processor.
[0024] The beneficial effects of this invention are as follows: This invention collects contact stage displacement signals, actual pin mark depth, and electrical test results. It utilizes Shannon entropy to quantify contact stability, normalized depth offset index, and continuous failure risk penalty factor to construct a composite working condition entropy vector, comprehensively characterizing the instantaneous state of the probe-pad contact interface, making working condition identification more sensitive and accurate. Based on this, a two-layer adaptive fuzzy inference mechanism is employed to generate a dynamic weight vector. In the event of severe anomalies, the weights are rapidly adjusted to ensure test safety. During fluctuations in normal working conditions, a continuously smooth multi-objective weight transition is achieved through an offline-trained neural fuzzy network. This ensures that the composite reward signal always matches the optimization priority of the current working condition, solving the optimization bias problem caused by a fixed weight reward function. This guides the strategy network to stably control the pin mark depth and maintain a high test pass rate throughout the probe's entire lifecycle.
[0025] Furthermore, the composite reward signal is decomposed into attribution vectors for each sub-objective through proportional attribution decomposition. These vectors are then used to weight the gradient terms of each sub-objective during policy gradient updates, ensuring that parameter updates focus on targets that are currently performing poorly or contributing significantly. This attribution-guided policy gradient mechanism accelerates the policy network's response to changing conditions across multiple objectives, effectively suppresses parameter drift caused by target masking, and improves the convergence stability of reinforcement learning during PCB testing fine-tuning, as well as its robustness to perturbations in board and probe states. Attached Figure Description
[0026] Figure 1 This is a flowchart illustrating the first embodiment of a reinforcement learning-based PCB test pin mark consistency optimization method according to the present invention.
[0027] Figure 2 This is a schematic diagram of the Hilbert amplitude envelope and contact stability entropy of the first embodiment of a reinforcement learning-based PCB test pin mark consistency optimization method of the present invention.
[0028] Figure 3 This is a schematic diagram of the material correction pin mark depth offset index of the first embodiment of the PCB test pin mark consistency optimization method based on reinforcement learning of the present invention.
[0029] Figure 4 This diagram illustrates the change in reward attribution during the probe wear stage, representing a first embodiment of a reinforcement learning-based PCB test pin mark consistency optimization method of the present invention. Detailed Implementation
[0030] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0031] Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] Example 1: As Figure 1 The diagram shown is a flowchart of the first embodiment of a PCB test pin mark consistency optimization method based on reinforcement learning according to the present invention. The first embodiment of the PCB test pin mark consistency optimization method based on reinforcement learning according to the present invention is presented.
[0033] In the first embodiment, the reinforcement learning-based PCB test pin mark consistency optimization method includes: Step S10: Obtain the multidimensional evaluation data of the previous acupuncture cycle, and perform the composite working condition characterization task based on the multidimensional evaluation data using entropy measurement and risk quantification mechanism, and output the composite working condition characterization vector. The multidimensional evaluation data from the previous needle-piercing cycle refers to the traceable data set generated by the flying probe testing equipment during the previous round of probe contact, pressure, holding, and electrical performance testing. This data includes at least high-frequency displacement sensor data, actual needle-piercing depth, pass / fail indicators for each test channel, and contact resistance transient recovery time. The entropy measurement and risk quantification mechanism does not simply read whether a test was successful. Instead, it first performs a Hilbert transform on the displacement signal within a preset time window to obtain an instantaneous amplitude sequence that reflects the contact vibration envelope. Then, it extracts peak points from this instantaneous amplitude sequence and divides these peak points into different amplitude intervals according to a preset number. The contact stability entropy is obtained based on the proportion of peak points in each interval to the total number of peak points. If no peak points are extracted within the preset time window, the contact stability entropy is set to a preset default entropy value to avoid interruptions in subsequent inference due to empty samples.
[0034] Within the same processing chain, the actual pin penetration depth is compared with the preset target pin mark depth. A preset material impact sensitivity correction factor corresponding to the current PCB material type is used to correct the sensitivity of different materials to overshoot or undershoot. The correction factor for ceramic substrates is greater than 1, ensuring that depth overshoot is more fully reflected; the correction factor for flexible substrates is less than 1, preventing excessive amplification of normal deviations caused by material elasticity. The pass / fail flags of each test channel, along with the contact resistance transient recovery time, jointly determine the failure risk factor: all channels pass with a recovery time less than 10 ms, indicating low risk; all channels pass but the recovery time reaches or exceeds 10 ms, indicating transient risk; and any channel fails, indicating high risk. The contact stability entropy, normalized pin mark depth offset index, and failure risk factor, combined according to a preset order, form a composite operating condition characterization vector.
[0035] The purpose of this composite working condition representation vector is to compress the original displacement waveform, pin mark depth, and electrical test results into a unified state expression that can be directly read by subsequent reinforcement learning. The contact stability entropy preserves the transient fluctuation of the probe-pad contact, the normalized pin mark depth offset index preserves the pin mark geometric consistency information, and the failure risk factor preserves the electrical test reliability information. When generating the dynamic weight vector in step S20, it is not necessary to re-examine the entire sensor waveform, nor is it necessary to make a rough judgment based solely on pass or fail flags. Therefore, the state recognition results can be stably transmitted to the reward structure adjustment stage.
[0036] Using common fixed threshold schemes as a reference, these schemes often only compare whether the contact resistance exceeds the threshold and whether the pin mark depth falls within the allowable range, failing to distinguish gradual changes caused by slight probe wear, micro-oxidation of the pad surface, or differences in board rigidity. Through this step, a decrease in contact stability will first manifest as a dispersion in the peak amplitude distribution, leading to an increase in entropy. Depth overshoot on the ceramic substrate will be amplified by the material correction factor, and tests with longer recovery times but not yet failed can also be recorded as transitional risks. Thus, subsequent strategy adjustments address a state characterization that includes degradation trends, rather than a failure alarm triggered after the fact.
[0037] On a batch testing line for hybrid PCB boards, the preset target needle mark depth for a certain ceramic substrate was 12 μm, and the actual depth in the previous needle mark cycle was 17 μm. The material impact sensitivity correction factor was set to 1.3. The peak points obtained after Hilbert transform of the displacement waveform were mainly concentrated in a few amplitude ranges, and the contact stability entropy could be recorded as 0.42. All test channels passed, and the contact resistance transient recovery time was 8 ms. Based on this, the system identified the depth offset as a medium offset amplified by the material, set the failure risk factor to low risk, and output a composite working condition characterization vector composed of a depth offset result of 0.42 (approximately 54.2%) and a low-risk flag, providing specific state basis for the next step of weighted inference.
[0038] Step S20: Based on the composite working condition representation vector, a hierarchical adaptive fuzzy inference mechanism is used to perform the dynamic weight vector generation task and output the dynamic weight vector; The hierarchical adaptive fuzzy inference mechanism refers to weighting the failure risk factor, normalized pinhole depth offset index, contact stability entropy, and contact resistance transient recovery time in the composite working condition representation vector according to two paths: emergency response and fine-grained balancing. The coarse-grained fuzzy rule response layer handles working conditions with significant risks or significant depth offsets: when the failure risk factor is greater than 0.7, the dynamic weights first tilt towards the failure risk target; when the failure risk does not exceed the threshold but the normalized pinhole depth offset index is greater than 80%, the dynamic weights preferentially tilt towards the pinhole depth target. These coarse-grained rules ensure that abnormal working conditions are not delayed by the complex inference process.
[0039] Within a typical fluctuation range where the failure risk factor is no greater than 0.7 and the depth offset index is no greater than 80%, the system invokes a pre-trained adaptive neural fuzzy inference network. This network is trained offline using historically optimal needle-pricking data. Gaussian membership functions are set for input variables such as contact stability entropy, depth offset, failure risk, and contact resistance transient recovery time. A fuzzy rule base is formed from historical expert-tuned data. The network learns a continuous mapping between contact state and fine-grained weights. For example, as the contact stability entropy gradually increases and the contact resistance recovery time approaches 15 ms from 10 ms, the weight corresponding to contact stability is smoothly increased while the weight of needle mark depth is decreased to avoid abrupt weight switching that could cause a jump in the reward signal.
[0040] After the dynamic weight vector is output, step S30 applies its three weight components to the contact stability score, pin mark depth score, and failure risk score, respectively. The technical effect of this step is to transform "what should be prioritized for optimization" into calculable reward structure parameters. When the failure risk suddenly increases, subsequent rewards will place greater emphasis on avoiding failure; when the pin mark depth shifts significantly, subsequent rewards will place greater emphasis on depth regression; when probe wear causes a decrease in contact stability but has not yet triggered a hard failure, the contact stability weight can be increased in advance, allowing the policy network to receive adjustment directions before failure occurs.
[0041] Fixed-weight multi-objective rewards remain constant throughout the probe's lifespan, making it easy for the strategy to continue following old optimization preferences even after probe state changes. For example, in the early stages, the focus might be on needle mark depth accuracy, but maintaining depth priority in the mid-to-late stages when contact recovery slows down would cause the strategy to ignore stable contact. This step combines coarse-grained rules with fine-grained neural fuzzy networks to enable rapid response to abnormal conditions and continuous transition to normal fluctuations. The weight changes have a physical state origin rather than being manually periodically tuned, thus more closely reflecting the actual operating conditions of the test line.
[0042] Near the 500th needle insertion, if the composite condition characterization vector shows a contact stability entropy of approximately 0.65, a depth offset of approximately 12%, a failure risk factor of 0.3, and a contact resistance transient recovery time of approximately 14 ms, then neither the failure risk nor the depth offset triggers the coarse-grained emergency rule. However, the contact stability and recovery time already indicate that the probe has entered the mid-stage of wear. The fine-grained inference network will increase the weight of the contact stability objective to a higher level and appropriately reduce the weight of the depth objective. In this way, subsequent composite rewards will no longer mistakenly treat "still good depth" as excellent overall performance, but will instead incorporate contact degradation into a more significant optimization direction.
[0043] Step S30: Based on the composite working condition representation vector and the dynamic weight vector, a weighted reward aggregation mechanism is used to perform the composite reward signal generation task, and the composite reward signal is output; The weighted reward aggregation mechanism involves first extracting contact stability entropy, normalized needle mark depth offset index, and failure risk factor from the composite working condition characterization vector, and then generating contact stability score, needle mark depth score, and failure risk score respectively. The contact stability score is negatively correlated with the contact stability entropy; a higher entropy value indicates more chaotic contact fluctuations, resulting in a lower score. The needle mark depth score is based on whether the actual needle depth falls within the target range of 5μm to 20μm and whether it is close to the preset target needle mark depth; the closer the depth is to the target median, the higher the score. The failure risk score is set based on the pass or fail results of each test channel, with a positive reward for pass and a negative penalty for failure.
[0044] After the three scores are generated, the system reads the weight components corresponding to contact stability, pinhole depth, and failure risk from the dynamic weight vector, adds each score to the total reward value according to its corresponding weight, and encapsulates the contact stability score, pinhole depth score, failure risk score, and total reward value into a composite reward signal. This output not only contains a total reward for reinforcement learning to evaluate the merits of the action, but also retains the scores of each sub-objective, thus allowing it to be further decomposed into a reward attribution vector in step S40.
[0045] This step truly integrates state representation and weight inference into the reinforcement learning feedback signal. If the contact stability entropy is high and step S20 has already increased the contact stability weight, even if the pinhole depth score is high, the composite reward will decrease due to the low stability score; if the depth shift is severe, the pinhole depth score will have a stronger impact on the total reward under high weight. In this way, the reward signal received by the policy network always corresponds to the pinhole quality dimension that most needs correction at the moment, rather than compressing all targets into the same reward ratio.
[0046] Fixed-weight rewards often suffer from sub-target masking: when the pinhole depth falls within the target range, the total reward may still be high even if contact stability has decreased, causing the policy network to misjudge the current action as valid. This step uses dynamic weights in reward aggregation to prevent depth advantages from masking contact degradation in the long run, and to prevent the risk of contact failure from being offset by better geometric pinhole results. The composite reward signal thus better reflects the true test quality and provides more targeted learning feedback for subsequent policy gradients.
[0047] After a needle prick, if the contact stability score is 0.25, the needle mark depth score is 0.95, and the failure risk score is 0.5, while the dynamic weight vector corresponds to contact stability 0.55, needle mark depth 0.20, and failure risk 0.25, the system multiplies each weight by its corresponding score and sums the results, yielding a total reward of approximately 0.4525. With fixed weights, the depth score might account for a larger proportion and award a higher reward. The lower total reward under dynamic weights indicates that while the current action maintained a good needle mark depth, it failed to address the contact stability issue, thus prompting subsequent strategies to be adjusted towards improving contact quality.
[0048] Step S40: Based on the composite reward signal and the dynamic weight vector, a proportional attribution decomposition mechanism is used to perform a reward attribution calculation task, and a reward attribution vector is output. The proportional attribution decomposition mechanism involves extracting contact stability score, pinhole depth score, and failure risk score from the composite reward signal, and multiplying each score by its corresponding weighted component in the dynamic weight vector to obtain three weighted reward contribution values. Subsequently, the system uses the sum of these three weighted contribution values as a normalization benchmark, and introduces a preset minimum positive number to avoid a zero denominator, converting each weighted contribution value into a contact stability attribution component, a pinhole depth attribution component, and a failure risk attribution component. These three components are combined in a fixed order to form the reward attribution vector.
[0049] The direct effect of this step is to decompose the total reward obtained in step S30 into a contribution structure with physical meaning. The total reward can only indicate the overall quality of the needle insertion action, but it cannot indicate whether the high or low reward is mainly due to contact stability, needle mark depth, or failure risk. The reward attribution vector retains the relative contribution of each type of sub-objective to the total reward. Step S50 can then use this vector to determine the gradient guidance weights of each sub-objective, enabling the policy network update to have directional interpretability.
[0050] Conventional policy gradient methods typically accept only a single scalar reward, with all parameter updates scaling along the same advantage estimation direction. This makes it difficult to distinguish whether a decrease in reward is due to deteriorating contact quality, excessive depth shift, or increased risk of test failure. Proportional attribution decomposition, however, assigns responsibility to the reward signal before it enters the policy update process. This reduces the possibility of high-molecular-weight targets obscuring low-molecular-weight targets and avoids applying the same update intensity to all action parameters. For multi-objective pinhole optimization scenarios, this attribution structure makes the learning process more likely to adjust to the actual defect sources.
[0051] Continuing with the aforementioned reward scenario, the total reward is approximately 0.4525. The three weighted contributions from contact stability, pinhole depth, and failure risk are 0.55 x 0.25, 0.20 x 0.95, and 0.25 x 0.5, respectively. After signed normalization, the three attribution components are approximately 0.304, 0.420, and 0.276, respectively. This result indicates that pinhole depth contributes the largest share to the reward, contact stability contributes a weaker but not eliminated contribution, and failure risk contributes a relatively small contribution. Subsequent strategy updates can thus maintain the advantage of depth control while continuing to compensate for insufficient contact stability.
[0052] Step S50: Based on the reward attribution vector, the reinforcement learning strategy network update task is executed using the attribution-guided gradient mechanism, and the PCB test pin mark consistency optimization parameters are output.
[0053] The attribution-guided strategy gradient mechanism refers to using the contact stability attribution component, needle mark depth attribution component, and failure risk attribution component in the reward attribution vector as gradient-guided weights for the three sub-objectives. The strategy network takes the strategy state, composed of a composite working condition representation vector, as input and outputs control parameters to adjust the probability distribution of actions for the next needle insertion cycle. The system then determines the advantage estimate of each sub-objective based on the deviation of the corresponding score value from the preset sub-objective baseline estimate. Subsequently, the advantage estimates of each sub-objective and their corresponding attribution components jointly contribute to the gradient calculation of the strategy network's objective function, ultimately updating the strategy network parameters in the direction that increases the objective function.
[0054] The output PCB test pin mark consistency optimization parameters can be understood as the control parameter adjustment results or the basis for generation used in the next pin-piercing cycle after the policy network update, including policy parameters related to control actions such as pressing speed, pressing depth fine-tuning, and contact maintenance time. This output follows the attribution structure of step S40, ensuring that the policy update is not merely an overall amplification or reduction based on the total reward, but rather adjusts the learning focus based on the contribution differences of the three sub-objectives: contact stability, pin mark depth, and failure risk. This makes the control actions in the next cycle more closely aligned with the current pin-piercing quality shortcomings. Furthermore, after the reinforcement learning policy network update task outputs the PCB test pin mark consistency optimization parameters, the control system performs closed-loop driving at the hardware level based on these optimization parameters. Specifically, the optimized parameters include at least: the adaptive deceleration look-ahead position of the Z-axis servo motor, the dynamic proportional-integral-derivative (PID) gain coefficient of the hardware force control closed loop, and the preload compensation displacement of the elastic element under the target contact torque; after receiving the optimized parameters, the motion control card directly rewrites the underlying servo loop register and trajectory interpolation counter before the start of the next needle puncture cycle, thereby achieving micron-level adaptive closed-loop control of the physical needle mark depth of the next needle puncture cycle by physically changing the current output waveform and deceleration curve of the motor.
[0055] Compared to fixed composite gradient updates, attribution-guided policy gradients can reduce the solidification of target preferences. When the probe enters the wear stage, fixed gradients may continue to update along the depth optimization direction, leading to a delayed response due to deterioration in contact stability. This step, when the reward attribution vector shows that contact stability is the main risk, enhances the update intensity of action parameters related to stable contact, causing the policy network to tend to reduce the pressing speed, extend the contact holding time, or choose a more conservative pressing adjustment magnitude. If the attribution structure shows that needle mark depth offset is dominant, the update focus shifts to depth control.
[0056] In continuous testing during the severe wear phase, if the reward attribution vectors for the most recent dozens of cycles show an average contact stability attribution component of approximately 0.65, while the pinhole depth and failure risk attribution components are low, the policy network, after collecting this batch of samples, will amplify the influence of contact stability-related sub-objectives in the gradient calculation. After several parameter updates, the output optimized parameters will favor lower pressure rates and longer contact maintenance times, causing the contact stability score to gradually recover. Once the contact state recovers, the reward attribution structure tends to be balanced, and the policy network will automatically return to an update direction that balances pinhole depth and test reliability.
[0057] To enable those skilled in the art to more clearly understand the intelligent control architecture of this invention, the online reinforcement learning (RL) closed-loop control loop constructed in this embodiment abstracts the probe testing mechanism, PCB board, and sensor network together as a physical environment; wherein: (1) The State Space is characterized by the composite working condition representation vector output in step S10. Its physical essence reflects the mechanical force state, deformation residual and electrical contact impedance of the probe tip at the end of the current needle puncture cycle. (2) The Action Space is characterized by the PCB test pin mark consistency optimization parameters output in step S50 and applied to the motion control card. Its physical essence is the feedforward correction behavior of the motor feed trajectory and servo stiffness in the next pin-piercing cycle. (3) The reward mechanism is constructed through fuzzy reasoning and proportional attribution decomposition in steps S20 to S40, providing quantitative causal guidance for the policy gradient evolution of the agent.
[0058] For example, such as Figure 2 As shown, during a specific pinning cycle in the batch testing of ceramic substrates, the contact displacement signal acquired by the high-frequency displacement sensor still indicates that the surface has completed the pressing and rebound processes. However, the instantaneous amplitude envelope obtained by Hilbert transform shows that the vibration peak distribution after the probe contacts the pad has changed from concentrated to dispersed. The system extracts the peak points from the instantaneous amplitude sequence and calculates the contact stability entropy after dividing the amplitude interval through K-means clustering, which can quantify this chaotic state of energy release at the contact interface in advance. The advantage of this technology is that the system does not have to wait until the test channel fails to respond. Instead, it can identify the decreasing trend of contact stability from the displacement waveform in the middle of probe wear and before the contact resistance recovery time is completely out of control, providing a more sensitive condition characterization for subsequent dynamic weight adjustment.
[0059] like Figure 3 As shown, even with the same actual pin depth shift from 12μm to 17μm, the risk implications for ceramic and flexible substrates are different. Ceramic substrates are more sensitive to impact and overshoot, with a material impact sensitivity correction factor greater than 1, thus amplifying the normalized pin depth shift index. Flexible substrates, on the other hand, have a certain degree of elasticity, with a correction factor less than 1, so the same depth shift is not over-penalized. In this way, the system can incorporate material differences into the composite condition representation vector, rather than using the same depth shift threshold for all PCB materials. The advantage of this technique is that the reinforcement learning strategy can obtain state inputs closer to actual damage risks when facing different material batches, reducing the possibility of underestimating overshoot on ceramic substrates or misjudging normal elastic fluctuations on flexible substrates as severe shifts.
[0060] like Figure 4As shown, the reward attribution structure changes significantly as the probe progresses from initial use to mid-wear and then to severe wear. Initially, the attribution for probe depth is relatively high, indicating that the strategy primarily focuses on maintaining depth consistency. During mid-wear, the attribution for contact stability gradually increases, suggesting that interface fluctuations become the primary concern. In the severe wear stage, if channel failure occurs, the attribution for failure risk rises rapidly. The advantage of this technique is that the reward attribution vector can shift the policy update focus according to the probe's state and testing risk, preventing the reinforcement learning process from becoming fixed in a single stage of optimization preference.
[0061] Example 2: Furthermore, the present invention provides a PCB test pin mark consistency optimization system based on reinforcement learning, which employs a PCB test pin mark consistency optimization method based on reinforcement learning in the above embodiments, and can solve the technical problem of PCB test pin mark consistency optimization based on reinforcement learning. The beneficial effects of the PCB test pin mark consistency optimization system based on reinforcement learning provided by the present invention are the same as those of the PCB test pin mark consistency optimization method based on reinforcement learning provided in the above embodiments, and other technical features of the PCB test pin mark consistency optimization system based on reinforcement learning are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0062] Example 3: This invention provides a PCB test pin trace consistency optimization device based on reinforcement learning. The device includes at least one processor and a memory communicatively connected to the processor. The memory stores instructions executable by the processor, which are then executed to enable the processor to perform the reinforcement learning-based PCB test pin trace consistency optimization method described in Example 1. This reinforcement learning-based PCB test pin trace consistency optimization device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. This reinforcement learning-based PCB test pin trace consistency optimization device is merely an example and should not limit the functionality or scope of this invention. A reinforcement learning-based PCB test pin conformance optimization device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes based on a program stored in a read-only memory or a program loaded from a storage device into a random access memory. The random access memory also stores various programs and data required for the operation of the reinforcement learning-based PCB test pin conformance optimization device. The processing unit, read-only memory, and random access memory are interconnected via a bus. An I / O interface is also connected to the bus. Typically, the following systems can be connected to the I / O interface: input devices including touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices including liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices including magnetic tapes, hard disks, etc.; and communication devices. The communication device allows the reinforcement learning-based PCB test pin conformance optimization device to communicate wirelessly or wiredly with other devices to exchange data. While a reinforcement learning-based PCB test pin conformance optimization device with various systems has been described, it should be understood that implementation or possession of all described systems is not required. It can be implemented alternatively or with more or fewer systems.
[0063] Example 4: This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the reinforcement learning-based PCB test pin mark consistency optimization method described above. The computer program product provided by this invention can solve the technical problem of PCB test pin mark consistency optimization based on reinforcement learning. Compared with the prior art, the beneficial effects of the computer program product provided by this invention are the same as those of the reinforcement learning-based PCB test pin mark consistency optimization method provided in the above embodiments, and will not be repeated here.
[0064] In particular, according to the embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a read-only memory. When the computer program is executed by a processing device, it performs the functions defined in the methods of the embodiments disclosed in this invention.
[0065] It should be understood that the various parts disclosed in this invention can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.
[0066] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the present invention and its equivalents, the present invention also intends to include these modifications and variations.
Claims
1. A reinforcement learning-based method for optimizing PCB test pin mark consistency, characterized in that, The methods include: Step S10: Obtain the multidimensional evaluation data of the previous acupuncture cycle, and perform the composite working condition characterization task based on the multidimensional evaluation data using entropy measurement and risk quantification mechanism, and output the composite working condition characterization vector. Step S20: Based on the composite working condition representation vector, a hierarchical adaptive fuzzy inference mechanism is used to perform the dynamic weight vector generation task and output the dynamic weight vector; Step S30: Based on the composite working condition representation vector and the dynamic weight vector, a weighted reward aggregation mechanism is used to perform the composite reward signal generation task, and the composite reward signal is output; Step S40: Based on the composite reward signal and the dynamic weight vector, a proportional attribution decomposition mechanism is used to perform a reward attribution calculation task, and a reward attribution vector is output. Step S50: Based on the reward attribution vector, the reinforcement learning strategy network update task is executed using the attribution-guided gradient mechanism, and the PCB test pin mark consistency optimization parameters are output.
2. The PCB test pin mark consistency optimization method based on reinforcement learning as described in claim 1, characterized in that, Step S10, which involves acquiring multidimensional evaluation data from the previous acupuncture cycle, performing a composite working condition characterization task based on the multidimensional evaluation data using entropy measurement and risk quantification mechanisms, and outputting a composite working condition characterization vector, specifically includes: Step S101: Read the high-frequency displacement sensor data of the previous acupuncture cycle from the multidimensional evaluation data; perform Hilbert transform on the displacement signal corresponding to the high-frequency displacement sensor data within a preset time window to obtain an instantaneous amplitude sequence; extract peak points from the instantaneous amplitude sequence and divide the peak points into a preset number of amplitude intervals using the K-means clustering method; calculate the information entropy according to the proportion of the number of peak points in each amplitude interval to the total number of peak points to obtain the contact stability entropy, wherein the contact stability entropy satisfies: in, This represents the contact stability entropy; This represents the amplitude range of the preset number; Indicates the amplitude interval index, and ; Indicates the first A range of amplitude values; Indicates falling into the first The number of peak points within each amplitude range; This indicates the total number of peak points; This represents the natural logarithm operation; when the total number of peak points is 0, the contact stability entropy is set to a preset default entropy value; Step S102: Read the actual needle depth of the previous needle-piercing cycle from the multidimensional evaluation data, obtain the preset material impact sensitivity correction factor corresponding to the current PCB material type, and calculate the normalized needle depth offset index based on the actual needle depth, the preset target needle mark depth, and the preset material impact sensitivity correction factor. The normalized needle depth offset index satisfies: in, This represents the normalized needle mark depth offset index; This represents the preset material impact sensitivity correction factor; This indicates the actual needle insertion depth; This indicates the preset target pin mark depth, and the preset target pin mark depth is a value greater than 0; when the current PCB material type is a ceramic substrate, the preset material impact sensitivity correction factor is greater than 1; when the current PCB material type is a flexible substrate, the preset material impact sensitivity correction factor is less than 1. Step S103: Read the pass / fail flag and contact resistance transient recovery time of each test channel from the multidimensional evaluation data, and determine the failure risk factor based on the pass / fail flag and the contact resistance transient recovery time; wherein, when all the test channels pass and the contact resistance transient recovery time is less than a preset recovery time threshold, the failure risk factor is set to 0; when all the test channels pass and the contact resistance transient recovery time is not less than the preset recovery time threshold, the failure risk factor is set to 0.3; when at least one of the test channels fails, the failure risk factor is set to 1; the preset recovery time threshold is 10 ms; Step S104: According to a preset arrangement order, combine the contact stability entropy, the normalized needle mark depth offset index, and the failure risk factor into the composite working condition characterization vector, wherein the composite working condition characterization vector satisfies: in, This represents the composite working condition characterization vector; Indicates the current moment of strategy decision-making; This refers to the failure risk factor.
3. The PCB test pin mark consistency optimization method based on reinforcement learning as described in claim 2, characterized in that, Step S20, which involves using a hierarchical adaptive fuzzy inference mechanism to generate a dynamic weight vector based on the composite working condition representation vector and outputting the dynamic weight vector, specifically includes: Step S201: Input the failure risk factor and the normalized pin mark depth offset index into a preset coarse-grained fuzzy rule response layer to obtain a coarse-grained weight vector; wherein, the coarse-grained weight vector includes three coarse-grained weight components arranged in the order of contact stability, pin mark depth, and failure risk; the preset risk threshold is 0.7, and the preset depth offset threshold is 80%; when the failure risk factor is greater than the preset risk threshold, the coarse-grained weight vector is set to... When the failure risk factor is not greater than the preset risk threshold, and the normalized needle mark depth offset index is greater than the preset depth offset threshold, the coarse-grained weight vector is set to... ; Step S202: When the failure risk factor is not greater than the preset risk threshold and the normalized pin mark depth offset index is not greater than the preset depth offset threshold, the contact stability entropy, the normalized pin mark depth offset index, the failure risk factor and the contact resistance transient recovery time are input into the pre-trained adaptive neural fuzzy inference network to obtain a fine-grained weight vector. Step S203: Determine the fusion coefficient based on the failure risk factor, and use the fusion coefficient to perform weighted fusion of the coarse-grained weight vector and the fine-grained weight vector to obtain the dynamic weight vector, wherein the dynamic weight vector satisfies: in, Represents the dynamic weight vector, and ; , and These represent the dynamic weight components corresponding to contact stability, pin mark depth, and failure risk, respectively. This represents the coarse-grained weight vector; This represents the fine-grained weight vector; This represents the fusion coefficient, and the fusion coefficient is equal to the failure risk factor.
4. The PCB test pin mark consistency optimization method based on reinforcement learning as described in claim 3, characterized in that, In step S202, the adaptive neural fuzzy inference network is obtained through offline training on historical best needle insertion data. The adaptive neural fuzzy inference network sets multiple Gaussian membership functions for each input variable and constructs a fuzzy rule base based on the historical best needle insertion data. The adaptive neural fuzzy inference network is used to learn the nonlinear continuous mapping relationship between the contact stability entropy, the contact resistance transient recovery time, and the fine-grained weight vector. When the contact stability entropy increases to 0.6 and the contact resistance transient recovery time increases to 15 ms, the adaptive neural fuzzy inference network continuously adjusts the fine-grained weight component corresponding to the contact stability from 0.3 to 0.5 and the fine-grained weight component corresponding to the needle mark depth from 0.5 to 0.
3.
5. The PCB test pin mark consistency optimization method based on reinforcement learning as described in claim 3, characterized in that, Step S30, which involves generating a composite reward signal based on the composite working condition representation vector and the dynamic weight vector using a weighted reward aggregation mechanism, and outputting the composite reward signal, specifically includes: Step S301: Calculate the contact stability score based on the contact stability entropy and the preset entropy-stability scoring function to obtain the contact stability score. The contact stability score is negatively correlated with the contact stability entropy. Step S302: Calculate the needle mark depth score based on the normalized needle mark depth offset index and the preset needle mark depth scoring rule to obtain the needle mark depth score. Wherein, when the actual needle insertion depth is within the preset target depth range of 5μm to 20μm, and the deviation between the actual needle insertion depth and the preset target needle mark depth decreases, the needle mark depth score increases. Step S303: Calculate the switch-type failure risk score based on the failure risk factor and the pass / fail flag to obtain the failure risk score. When all the test channels pass, the failure risk score is set to a preset positive reward value; when at least one of the test channels fails, the failure risk score is set to a preset negative reward value. Step S304: Using the three dynamic weight components in the dynamic weight vector, the contact stability score, the pin mark depth score, and the failure risk score are weighted and summed to obtain the total reward value. The contact stability score, pin mark depth score, failure risk score, and total reward value are then encapsulated into the composite reward signal, wherein the total reward value satisfies: in, This represents the total reward value.
6. The PCB test pin mark consistency optimization method based on reinforcement learning as described in claim 5, characterized in that, Step S40, which involves performing a reward attribution calculation task based on the composite reward signal and the dynamic weight vector using a proportional attribution decomposition mechanism, and outputting a reward attribution vector, specifically includes: Step S401: Extract the contact stability score, the pin mark depth score, and the failure risk score from the composite reward signal, and multiply the contact stability score, the pin mark depth score, and the failure risk score by the corresponding dynamic weight components in the dynamic weight vector to obtain three weighted reward contribution values. Step S402: Perform signed normalization processing on the three weighted reward contribution values to obtain the contact stability attribution component, the pin mark depth attribution component, and the failure risk attribution component, wherein the attribution components satisfy: in, Indicates the first Each attribution component; Indicates the scoring item number, and when When corresponding to the contact stability score, when When corresponding to the needle mark depth score, when The corresponding failure risk score; Indicates the first The rating value corresponding to each rating item; Indicates the first The rating value corresponding to each rating item; Indicates the summation sequence number; This represents the minimum positive number preset to avoid a denominator of 0; and These represent the i-th and j-th dynamic weight components, respectively. Step S403: Combine the contact stability attribution component, the pin mark depth attribution component, and the failure risk attribution component in the order of arrangement to obtain the reward attribution vector.
7. The PCB test pin mark consistency optimization method based on reinforcement learning as described in claim 6, characterized in that, Step S50, which involves executing the reinforcement learning policy network update task based on the reward attribution vector using an attribution-guided gradient mechanism and outputting PCB test pin mark consistency optimization parameters, specifically includes: Step S501: Determine the contact stability attribution component, pinhole depth attribution component, and failure risk attribution component in the reward attribution vector as gradient-guided weights for the contact stability sub-objective, pinhole depth sub-objective, and failure risk sub-objective, respectively. Step S502: For each sub-objective, determine the sub-objective advantage estimate based on the deviation of the corresponding score value from the preset sub-objective baseline estimate, and calculate the gradient of the policy network objective function with respect to the policy network parameters according to the following formula: in, Represents the objective function of the policy network; Indicates the policy network parameters; This represents the strategy state composed of the composite operating condition representation vector; This indicates the adjustment action of the control parameters output by the strategy network for the next injection cycle; This represents the conditional probability that the policy network will output the control parameter adjustment action under the policy state; Indicates the first Sub-objective advantage estimation for each sub-objective; This represents the gradient operation applied to the network parameters of the policy; This represents the expected value calculation at the moment of strategy decision-making; Step S503: Along the direction that increases the objective function of the policy network, update the policy network parameters according to the gradient of the objective function of the policy network relative to the policy network parameters to obtain the PCB test pin mark consistency optimization parameters.
8. A reinforcement learning-based PCB test pin mark consistency optimization system, applied to the reinforcement learning-based PCB test pin mark consistency optimization method according to any one of claims 1 to 7, characterized in that, The system includes: The composite working condition characterization module is used to acquire multi-dimensional evaluation data of the previous needle puncture cycle, and to perform the composite working condition characterization task based on the multi-dimensional evaluation data using an entropy measurement and risk quantification mechanism, and output the composite working condition characterization vector. The dynamic weight generation module is used to perform a dynamic weight vector generation task based on the composite working condition representation vector using a hierarchical adaptive fuzzy inference mechanism, and outputs a dynamic weight vector. The composite reward generation module is used to perform a composite reward signal generation task based on the composite working condition representation vector and the dynamic weight vector using a weighted reward aggregation mechanism, and output a composite reward signal. The reward attribution decomposition module is used to perform a reward attribution calculation task based on the composite reward signal and the dynamic weight vector using a proportional attribution decomposition mechanism, and output a reward attribution vector. The policy network update module is used to perform a reinforcement learning policy network update task based on the reward attribution vector using an attribution-guided policy gradient mechanism, and output PCB test pin mark consistency optimization parameters.
9. A PCB test pin mark consistency optimization device based on reinforcement learning, characterized in that, The reinforcement learning-based PCB test pin mark consistency optimization device includes: a memory, a processor, and a reinforcement learning-based PCB test pin mark consistency optimization program stored in the memory and executable on the processor. When the reinforcement learning-based PCB test pin mark consistency optimization program is executed by the processor, it implements a reinforcement learning-based PCB test pin mark consistency optimization method according to any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a reinforcement learning-based PCB test pin mark consistency optimization program, which, when executed by a processor, implements a reinforcement learning-based PCB test pin mark consistency optimization method according to any one of claims 1 to 7.