A method and system for quantitative evaluation of valve cracks of a drilling pump

CN122567216BActive Publication Date: 2026-09-22NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611058593.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-16
Publication Date
2026-09-22
Estimated Expiration
2046-07-16

AI Technical Summary

Technical Problem

[0005]本申请提供一种钻井泵阀裂纹定量评估方法及系统,旨在解决现有技术在钻井泵阀裂纹评估中,阀裂纹损伤特征表现隐蔽、运行工况变化对监测数据影响较大以及裂纹严重程度难以实现稳定的细粒度定量评估的问题

Benefits of technology

本申请基于对现有技术问题的进一步分析和研究,认识到现有技术在钻井泵阀裂纹评估中,阀裂纹损伤特征表现隐蔽、运行工况变化对监测数据影响较大以及裂纹严重程度难以实现稳定的细粒度定量评估的问题,通过获取钻井泵运行过程中的振动信号和声学信号,并根据待评估阀所在泵缸确定目标振动信号,使裂纹评估过程能够同时获得与待评估阀位置相关的局部机械响应信息以及钻井泵整体运行产生的声学响应信息;进一步通过窗口化处理将连续运行信号转换为包括振动分量和声学分量的双模态信号窗口,使隐蔽、渐进的裂纹损伤特征能够在固定时间片段内被提取和比较;再通过参数不共享的时序特征编码网络分别生成振动特征词元序列和声学特征词元序列,并结合模态标识信息和时序位置信息构建双流特征词元序列,使振动模态和声学模态在融合前分别保持各自的物理来源特征和时序变化特征,避免异构监测数据被简单混合而削弱裂纹相关表征;同时,通过在裂纹评估编码模型中构建包括不同编码阶段历史表征的历史表征池,并分别基于振动模态混合查询和声学模态混合查询检索与相应模态匹配的历史表征,使不同模态能够按照各自的特征演化规律复用历史信息,并将检索到的历史表征作为残差表征参与注意力编码,由此得到兼具局部冲击信息、全局声学信息和多阶段历史表征信息的融合裂纹表征;最终根据融合裂纹表征生成钻井泵阀裂纹严重程度预测值并输出定量评估结果,从而能够提高阀裂纹隐蔽损伤特征的表征能力,降低运行工况变化对评估结果稳定性的影响,实现对钻井泵阀裂纹严重程度的细粒度、稳定定量评估。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122567216B_ABST
    Figure CN122567216B_ABST
Patent Text Reader

Abstract

The application discloses a kind of drilling pump valve crack quantitative evaluation method and system.The method obtains vibration signal and acoustic signal in the operation process of drilling pump, determines the target vibration signal corresponding to the valve to be evaluated and constitutes double-mode evaluation signal;Windowing processing obtains the double-mode signal window containing vibration component and acoustic component;Respectively through parameter non-sharing time sequence characteristic coding network generates vibration feature token sequence and acoustic feature token sequence, and constructs double-flow feature token sequence in combination with modal identification information and time sequence position information;History representation pool is constructed in crack evaluation coding model, based on vibration mode mixed query and acoustic mode mixed query, match history representation is retrieved and attention residual coding is carried out, to obtain fusion crack representation, and then output valve crack severity prediction value.Thereby improve the hidden crack damage feature representation capability and the quantitative evaluation stability under different working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of drilling pump condition monitoring technology, and in particular to a method and system for quantitative assessment of drilling pump valve cracks. Background Technology

[0002] Drilling pumps are crucial power equipment in oil drilling mud circulation systems, primarily used to deliver high-pressure drilling fluid to the bottom of the well to maintain fluid circulation and clean the wellbore. During operation, valve components such as the suction and discharge valves on the hydraulic end of the drilling pump frequently open and close, and are subjected to long-term effects from high-pressure fluid impact, alternating loads, valve body collisions, fluid pulsation, and structural vibrations. These factors can easily lead to fatigue damage, decreased sealing performance, and crack propagation. Once cracks occur in the valve components, it can cause abnormal fit between the valve body and seat, affecting pump pressure stability and drilling fluid delivery efficiency. In severe cases, it can even cause abnormal equipment shutdown, impacting the continuity and safety of drilling operations.

[0003] Currently, the monitoring and diagnosis of drilling pump operation status typically involves collecting operational status data such as vibration, pressure, pump surge, temperature, or sound, and combining this data with signal processing, feature extraction, fault identification models, or empirical threshold judgments to identify issues such as pump valve leakage, hydraulic end anomalies, transmission component wear, or bearing failures. While these methods can reflect whether a drilling pump is in an abnormal operating state to some extent, in practical applications, drilling pump operating conditions are complex. Factors such as speed, load, medium flow conditions, and on-site noise can all affect the stability of monitoring data. Furthermore, the impact of valve cracks on equipment operation status at different stages of expansion is gradual and insidious; early or moderate cracks may not necessarily manifest as obvious fault phenomena. Therefore, existing monitoring methods are often more suitable for anomaly identification or fault type judgment, but they still lack the fine-grained characterization of valve crack severity, continuous assessment, and stability judgment under different operating conditions.

[0004] Therefore, in the assessment of cracks in drilling pump valves, the hidden characteristics of valve crack damage, the significant impact of changes in operating conditions on monitoring data, and the difficulty in achieving stable, fine-grained quantitative assessment of crack severity have become urgent problems to be solved. Summary of the Invention

[0005] This application provides a method and system for quantitative assessment of cracks in drilling pump valves, aiming to solve the problems in existing technologies for assessing cracks in drilling pump valves, such as the concealed nature of valve crack damage characteristics, the significant impact of changes in operating conditions on monitoring data, and the difficulty in achieving stable, fine-grained quantitative assessment of crack severity.

[0006] In a first aspect, a method for quantitative assessment of cracks in drilling pump valves, the method comprising: Vibration and acoustic signals during the operation of the drilling pump are acquired. The target vibration signal is determined from the vibration signals based on the pump cylinder where the valve to be evaluated is located. The target vibration signal and the acoustic signal are then combined to form a dual-mode evaluation signal. The dual-modal evaluation signal is windowed to obtain a dual-modal signal window, which includes vibration and acoustic components. The vibration component and the acoustic component are respectively encoded using a time-series feature coding network with non-shared parameters to obtain vibration feature word sequences and acoustic feature word sequences; A dual-stream feature word sequence is constructed based on the vibration feature word sequence, the acoustic feature word sequence, modal identification information, and temporal position information. The dual-stream feature word sequence is input into the crack evaluation coding model, and a historical representation pool is constructed in the crack evaluation coding model. The historical representation pool includes historical representations generated by the crack evaluation coding model at different coding stages. A vibration modal hybrid query is constructed for the vibration feature lexical sequence, and an acoustic modal hybrid query is constructed for the acoustic feature lexical sequence; wherein, the vibration modal hybrid query includes a learnable vibration query and an input-related vibration query generated based on the vibration feature lexical sequence, and the acoustic modal hybrid query includes a learnable acoustic query and an input-related acoustic query generated based on the acoustic feature lexical sequence; Based on the vibration mode hybrid query and the acoustic mode hybrid query, historical representations matching the vibration mode and historical representations matching the acoustic mode are retrieved from the historical representation pool, respectively. The retrieved historical representations are used as residual representations to participate in the attention encoding of the dual-stream feature word sequence to obtain the fused crack representation. Based on the fusion crack characterization, a predicted value for the severity of drilling pump valve cracks is generated, and a quantitative assessment result of drilling pump valve cracks is output based on the predicted value for the severity of drilling pump valve cracks.

[0007] Optionally, in the above scheme, acquiring vibration and acoustic signals during the operation of the drilling pump, determining the target vibration signal from the vibration signals based on the pump cylinder where the valve to be evaluated is located, and constructing a dual-mode evaluation signal from the target vibration signal and the acoustic signal, includes: Acquire multi-channel vibration and acoustic signals during the operation of the drilling pump to obtain operational status data; Obtain the pump cylinder position information corresponding to the valve to be evaluated, and determine the vibration channel of the corresponding pump cylinder from the multi-channel vibration signal based on the pump cylinder position information to obtain the target vibration channel; The vibration signal corresponding to the target vibration channel is extracted from the data collected in the operating state to obtain the target vibration signal; The target vibration signal and the acoustic signal are time-aligned and amplitude-preprocessed to obtain a synchronous dual-mode signal; The dual-mode evaluation signal is constructed based on the synchronous dual-mode signal.

[0008] Optionally, in the above scheme, the step of windowing the bimodal evaluation signal to obtain a bimodal signal window includes: The target vibration signal in the dual-modal evaluation signal is segmented according to the preset window length and preset window step size to obtain a vibration window sequence; The acoustic signal in the dual-modal evaluation signal is segmented according to the same window position as the target vibration signal to obtain an acoustic window sequence; The vibration window and acoustic window corresponding to the same window position are paired to obtain the dual-mode signal window.

[0009] Optionally, in the above scheme, the step of performing feature encoding on the vibration component and the acoustic component respectively through a time-series feature encoding network with non-shared parameters to obtain a vibration feature word sequence and an acoustic feature word sequence includes: The vibration components are input into a first temporal feature encoding network for local temporal feature extraction to obtain a vibration local feature sequence. The vibration local feature sequence is subjected to feature aggregation and word mapping to obtain the vibration feature word sequence; The acoustic components are input into a second temporal feature encoding network for local temporal feature extraction to obtain an acoustic local feature sequence; wherein, the network parameters of the first temporal feature encoding network and the second temporal feature encoding network are not shared; The acoustic local feature sequence is subjected to feature aggregation and word mapping to obtain the acoustic feature word sequence.

[0010] Optionally, in the above scheme, constructing a dual-stream feature word sequence based on the vibration feature word sequence, the acoustic feature word sequence, modal identification information, and temporal position information includes: Vibration mode identification information is generated based on the vibration mode, and acoustic mode identification information is generated based on the acoustic mode; Based on the arrangement of each feature word in the vibration feature word sequence and the acoustic feature word sequence, temporal position information is generated; The vibration mode identification information and the corresponding temporal position information are added to the vibration feature word sequence to obtain the vibration enhancement word sequence; The acoustic modality identifier information and the corresponding temporal position information are added to the acoustic feature word sequence to obtain the acoustic enhancement word sequence; The vibration-enhancing lexical sequence and the acoustic-enhancing lexical sequence are concatenated according to a preset lexical arrangement order to obtain the dual-stream feature lexical sequence.

[0011] Optionally, in the above scheme, the step of inputting the dual-stream feature lexical sequence into the crack evaluation coding model and constructing a historical representation pool in the crack evaluation coding model includes: The dual-stream feature word sequence is added to the historical representation pool as the initial historical representation to obtain the initial historical representation pool; In the current encoding stage of the crack evaluation encoding model, attention encoding is performed based on the current input representation to obtain a self-attention output representation; The self-attention output representation is added to the initial historical representation pool to obtain the first updated historical representation pool; Based on the self-attention output representation, a feedforward transformation is performed to obtain the feedforward output representation; The feedforward output representation is added to the first updated history representation pool to obtain the second updated history representation pool, and the second updated history representation pool is used as the history representation pool for the subsequent encoding stage.

[0012] Optionally, in the above scheme, constructing a vibration modal hybrid query for the vibration feature lexical sequence and constructing an acoustic modal hybrid query for the acoustic feature lexical sequence includes: Input-related features are extracted from the vibration feature word sequence to obtain input-related vibration queries; Obtain the learnable vibration query corresponding to the vibration mode, and perform a weighted combination of the input related vibration query and the learnable vibration query according to the vibration query gating parameter to obtain the vibration mode hybrid query; Input-related features are extracted from the acoustic feature word sequence to obtain input-related acoustic queries; A learnable acoustic query corresponding to an acoustic mode is obtained, and the input-related acoustic query and the learnable acoustic query are weighted and combined according to the acoustic query gating parameters to obtain the acoustic mode hybrid query; wherein, the learnable vibration query, the input-related vibration query, and the vibration query gating parameters correspond to different query parameters as the learnable acoustic query, the input-related acoustic query, and the acoustic query gating parameters, respectively.

[0013] Optionally, in the above scheme, based on the vibration mode hybrid query and the acoustic mode hybrid query, historical representations matching the vibration mode and historical representations matching the acoustic mode are retrieved from the historical representation pool, respectively, and the retrieved historical representations are used as residual representations to participate in the attention encoding of the dual-stream feature word sequence to obtain the fused crack representation, including: Construct a term-level query sequence based on the vibration mode mixed query and the acoustic mode mixed query; The historical representations in the historical representation pool are normalized to obtain historical key representations. The first historical retrieval weight corresponding to each feature word is determined based on the word-level query sequence and the historical key representation. The historical representations in the historical representation pool are weighted and aggregated according to the first historical retrieval weight to obtain the first historical residual representation. The first historical residual representation is used as a residual representation to participate in the multi-head self-attention encoding of the current input representation to obtain the self-attention output representation; The self-attention output representation is added to the historical representation pool to obtain the updated historical representation pool; Based on the term-level query sequence, a second historical retrieval is performed from the updated historical representation pool to obtain a second historical residual representation; The second historical residual representation is used as a residual representation to participate in the feedforward transformation of the self-attention output representation to obtain the encoded output representation; The encoded output representation is normalized and pooled to obtain the fusion crack representation.

[0014] Optionally, in the above scheme, the step of generating a predicted value for the severity of drilling pump valve cracks based on the fusion crack characterization, and outputting a quantitative assessment result of drilling pump valve cracks based on the predicted value for the severity of drilling pump valve cracks, includes: The fused crack characterization is input into the ordered prediction branch to obtain an estimated value of the ordered crack severity. The fused crack characterization is input into the regression prediction branch to obtain an estimate of the severity of the continuous crack. The severity estimates of ordered cracks and continuous cracks are fused together to obtain the predicted severity values ​​of the drilling pump valve cracks. The crack length assessment result and / or crack grade assessment result are determined based on the predicted crack severity value of the drilling pump valve; Based on the crack length assessment results and / or the crack grade assessment results, output the quantitative assessment results of the drilling pump valve cracks.

[0015] Secondly, a quantitative assessment system for cracks in drilling pump valves, the system comprising: The signal acquisition module is used to acquire vibration signals and acoustic signals during the operation of the drilling pump, determine the target vibration signal from the vibration signals according to the pump cylinder where the valve to be evaluated is located, and form a dual-mode evaluation signal by combining the target vibration signal and the acoustic signal. A window construction module is used to perform windowing processing on the dual-modal evaluation signal to obtain a dual-modal signal window, wherein the dual-modal signal window includes vibration components and acoustic components; The dual-stream coding module is used to encode the vibration component and the acoustic component separately through a time-series feature coding network with non-shared parameters, to obtain a vibration feature word sequence and an acoustic feature word sequence; The lexical construction module is used to construct a dual-stream feature lexical sequence based on the vibration feature lexical sequence, the acoustic feature lexical sequence, modal identification information, and temporal position information; The historical representation construction module is used to input the dual-stream feature word sequence into the crack evaluation coding model and construct a historical representation pool in the crack evaluation coding model. The historical representation pool includes historical representations generated by the crack evaluation coding model at different coding stages. A modal query construction module is used to construct a vibration modal hybrid query for the vibration feature lexical sequence and an acoustic modal hybrid query for the acoustic feature lexical sequence; wherein, the vibration modal hybrid query includes a learnable vibration query and an input-related vibration query generated based on the vibration feature lexical sequence, and the acoustic modal hybrid query includes a learnable acoustic query and an input-related acoustic query generated based on the acoustic feature lexical sequence; The attention residual encoding module is used to retrieve historical representations that match the vibration mode and the acoustic mode respectively from the historical representation pool based on the vibration mode mixed query and the acoustic mode mixed query, and to use the retrieved historical representations as residual representations to participate in the attention encoding of the dual-stream feature word sequence to obtain the fused crack representation; The crack assessment output module is used to generate a predicted value of the crack severity of the drilling pump valve based on the fused crack characterization, and output a quantitative assessment result of the drilling pump valve crack based on the predicted value of the crack severity of the drilling pump valve.

[0016] Compared with the prior art, this application has at least the following beneficial effects: Based on further analysis and research of existing technical problems, this application recognizes that existing technologies for assessing cracks in drilling pump valves suffer from several issues: concealed crack damage characteristics, significant impact of changes in operating conditions on monitoring data, and difficulty in achieving stable, fine-grained quantitative assessment of crack severity. This application addresses these problems by acquiring vibration and acoustic signals during drilling pump operation and determining the target vibration signal based on the pump cylinder where the valve to be assessed is located. This allows the crack assessment process to simultaneously obtain local mechanical response information related to the valve's location and acoustic response information generated by the overall operation of the drilling pump. Furthermore, windowing processing converts the continuous operating signal into a dual-modal signal window including vibration and acoustic components, enabling the extraction and comparison of concealed, progressive crack damage characteristics within a fixed time segment. Finally, a non-parameter-sharing temporal feature coding network generates vibration and acoustic feature word sequences respectively, and combines modal identification information and temporal position information to construct a dual-stream feature word sequence. This approach ensures that vibration and acoustic modes retain their respective physical origin and temporal variation characteristics before fusion, preventing the simple mixing of heterogeneous monitoring data from weakening crack-related characterization. Simultaneously, by constructing a historical characterization pool encompassing historical characterizations from different coding stages within the crack assessment coding model, and retrieving historical characterizations matching the corresponding modes based on vibration and acoustic mode mixed queries, different modes can reuse historical information according to their respective feature evolution laws. The retrieved historical characterizations are then used as residual characterizations in attention coding, resulting in a fused crack characterization that incorporates local impact information, global acoustic information, and multi-stage historical characterization information. Finally, based on the fused crack characterization, a predicted value for the severity of drilling pump and valve cracks is generated, and a quantitative assessment result is output. This improves the characterization capability of hidden damage characteristics of valve cracks, reduces the impact of changes in operating conditions on the stability of the assessment results, and achieves fine-grained, stable, and quantitative assessment of the severity of drilling pump and valve cracks. Attached Figure Description

[0017] Figure 1 A flowchart illustrating a method for quantitative assessment of cracks in drilling pump valves provided in one embodiment of this application; Figure 2 A schematic diagram of the overall framework of a method and system for quantitative assessment of cracks in drilling pump valves provided in one embodiment of this application; Figure 3 This is a schematic diagram of the structure of a modality-specific hybrid query attention residual coding block provided in one embodiment of this application; Figure 4 A schematic diagram of the structure of a drilling pump vibration-acoustic test bench provided in one embodiment of this application; Figure 5 This is a schematic diagram of the time-domain waveforms of vibration and acoustic signals under different crack severity levels, provided as an embodiment of this application. Figure 6A scatter plot of the relationship between the predicted severity of a drilling pump valve crack and the actual crack length, provided in one embodiment of this application. Figure 7 A schematic diagram of the normalized confusion matrix of the crack level prediction result of a drilling pump valve provided in one embodiment of this application; Figure 8 This is a schematic diagram illustrating the distribution of predicted and actual values ​​at different operating speeds, provided in one embodiment of this application. Figure 9 This is a schematic diagram of a modal-level inter-lexical attention matrix provided in one embodiment of this application; Figure 10 This is a schematic diagram of the attention residual weights for modal separation provided in one embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the technical features in the following embodiments can be combined with each other. The following embodiments are only used to explain the technical content of this application and are not intended to limit the scope of protection of this application.

[0019] This application provides a method and system for quantitative assessment of cracks in drilling pump valves, such as... Figure 1 As shown, this method takes the vibration and acoustic signals during the operation of the drilling pump as inputs and the crack length, crack grade, or crack severity of the drilling pump valve as outputs. It is suitable for crack damage assessment of valve components such as hydraulic end suction valves and discharge valves of drilling pumps, and can also be extended to other mechanical component damage assessment tasks with local structural impact and global acoustic radiation response.

[0020] like Figure 2 As shown, the quantitative assessment method for drilling pump valve cracks may include signal acquisition, window construction, dual-stream feature lexical encoding, modal identification and temporal location enhancement, historical characterization pool construction, modal hybrid query construction, attention residual encoding, and crack severity prediction output, among other processing steps. Figure 2 The vibration lexicon encoder and acoustic lexicon encoder in the model correspond to time-series feature coding networks with non-shared parameters, respectively. The modality-specific hybrid query attention residual coding block is used to fuse and encode the dual-stream feature lexicon sequences. The ordered prediction branch and regression prediction branch are used to output the final crack severity prediction value based on the fused crack characterization. Figure 3 The internal structure of the modality-specific hybrid query attention residual coding block is shown. This coding block may include structures such as historical representation pool, pre-attention historical retrieval, multi-head self-attention, pre-feedforward historical retrieval, and feedforward transformation.

[0021] like Figure 1As shown in the figure, this embodiment provides a method for quantitative assessment of cracks in drilling pump valves, which includes the following steps.

[0022] First, vibration and acoustic signals from the drilling pump during operation are acquired. Based on the location of the valve to be evaluated in the pump cylinder, the target vibration signal is determined from the vibration signals, and the target vibration signal and acoustic signal are combined to form a dual-modal evaluation signal. Vibration signals can be acquired by accelerometers installed in the drilling pump cylinder, hydraulic end housing, or near the valve assembly. Acoustic signals can be acquired by sound pressure sensors or microphones placed near the drilling pump. Vibration signals primarily reflect the valve assembly's opening and closing impact, valve body and seat collision, local structural vibration, and pump cylinder structural response. Acoustic signals primarily reflect the overall sound radiation, fluid disturbance, propagation path coupling, and changes in the ambient sound field during drilling pump operation. If the valve to be evaluated is located in the third pump cylinder, the vibration signal in the corresponding vibration channel of the third pump cylinder can be used as the target vibration signal, and combined with the global acoustic channel signal to form a dual-modal evaluation signal.

[0023] Secondly, the dual-modal evaluation signal is windowed to obtain a dual-modal signal window, which includes both vibration and acoustic components. Windowing processing can include filtering, mean removal, normalization, time synchronization, and fixed-length segmentation. Specifically, the target vibration signal can be segmented according to a preset window length, and the acoustic signal can be segmented according to the same window position, ensuring that each dual-modal signal window contains both the time-corresponding vibration and acoustic components. Figure 5 As shown, vibration signals and acoustic signals with different crack severity will exhibit differences in impact amplitude, periodic disturbance and acoustic response in the time domain waveform. Windowing processing can convert continuous running signals into sample units suitable for model input.

[0024] Then, the vibration and acoustic components are respectively encoded using a time-series feature coding network with non-shared parameters, resulting in vibration feature word sequences and acoustic feature word sequences. The time-series feature coding network can employ a one-dimensional convolutional network, a one-dimensional residual network, a lightweight residual network, a depthwise separable convolutional network, or other network structures suitable for local feature extraction of time-series signals. The vibration and acoustic components are fed into different time-series feature coding networks, with no parameter sharing between the two networks. This allows the vibration and acoustic modes to learn local time-series features that conform to their own physical sources, noise characteristics, and feature evolution laws.

[0025] Next, a dual-stream feature word sequence is constructed based on the vibration feature word sequence, the acoustic feature word sequence, modal identifier information, and temporal position information. Modal identifier information identifies whether the feature word originates from a vibration mode or an acoustic mode, while temporal position information identifies the position of the feature word within the corresponding modal sequence. Specifically, vibration modal embedding and positional embedding can be added to the vibration feature word sequence, and acoustic modal embedding and positional embedding can be added to the acoustic feature word sequence. Then, the sequences are concatenated according to a preset word order, for example, concatenating the vibration feature word sequence first, followed by the acoustic feature word sequence, to obtain the dual-stream feature word sequence.

[0026] Furthermore, the two-stream feature lexical sequence is input into the crack evaluation coding model, and a historical representation pool is constructed within the crack evaluation coding model. The historical representation pool includes historical representations generated by the crack evaluation coding model at different coding stages. Historical representations can include the initial two-stream feature lexical sequence, the preceding self-attention output representation, the preceding feedforward output representation, and the output representations from the preceding coding stages. The historical representation pool is used to store representations of different depths and semantic levels, enabling the current coding stage to adaptively select the required historical information from the historical representation pool.

[0027] Furthermore, a hybrid vibration modality query is constructed for the vibration feature lexical sequence, and a hybrid acoustic modality query is constructed for the acoustic feature lexical sequence. The hybrid vibration modality query includes a learnable vibration query and an input-related vibration query generated based on the vibration feature lexical sequence; the hybrid acoustic modality query includes a learnable acoustic query and an input-related acoustic query generated based on the acoustic feature lexical sequence. The learnable query represents the global historical retrieval preference formed by the corresponding modality during training, while the input-related query represents the historical retrieval needs of the current sample and the current crack state. The hybrid vibration modality query and the hybrid acoustic modality query employ different query parameters and different gating parameters to enable the two types of modalities to form different historical retrieval strategies.

[0028] Furthermore, based on vibration mode hybrid queries and acoustic mode hybrid queries, historical representations matching vibration modes and acoustic modes are retrieved from the historical representation pool, respectively. These retrieved historical representations are then used as residual representations in the attention encoding of the dual-stream feature term sequence to obtain a fused crack representation. Specifically, historical retrieval weights can be determined based on the similarity between the modal hybrid query and each historical representation in the historical representation pool. Then, historical representations in the historical representation pool are weighted and aggregated according to these historical retrieval weights to obtain historical residual representations. These historical residual representations, as residual paths, participate in the multi-head self-attention and feedforward transformation in the current encoding stage, enabling the current feature terms to perform information interaction both within and between modes, and to reuse shallow local impact features and deep global working condition features.

[0029] Finally, based on the crack characterization, a predicted value for the severity of drilling pump valve cracks is generated, and a quantitative assessment result of the drilling pump valve cracks is output based on this predicted value. The predicted value for the severity of drilling pump valve cracks can be expressed as crack length, crack grade, damage factor, percentage of remaining life, or maintenance risk level. In one specific embodiment, the predicted value for the severity of drilling pump valve cracks represents the valve crack length, with an output range of 0 to 10 mm, where 0 mm represents a healthy valve state, and 1 mm to 10 mm represent different crack lengths. Figure 6 As shown, the continuous prediction capability can be verified by the scatter distribution between the model's predicted values ​​and the actual crack lengths; for example... Figure 7 As shown, the crack severity discrimination capability can be verified using a normalized confusion matrix; as Figure 8 As shown, the stability of the model at different operating speeds can be verified by comparing the distribution of predicted and actual values ​​at different operating speeds.

[0030] In this embodiment, a dual-modal evaluation signal is constructed using the target vibration signal and the acoustic signal, enabling crack evaluation to simultaneously utilize the local mechanical response of the valve assembly and the global acoustic response of the drilling pump. Vibration feature word sequences and acoustic feature word sequences are obtained through a time-series feature encoding network with non-shared parameters, ensuring that the two heterogeneous modes retain their respective feature expressions before deep fusion. Attention residual encoding is performed through a historical representation pool and modal hybrid query, enabling different modes to retrieve different historical representations based on their own feature evolution laws. By fusing crack representations, a crack severity prediction value is output, enabling stable fine-grained quantitative evaluation even when crack features are hidden and operating conditions change.

[0031] like Figure 4 As shown, in one embodiment, quantitative assessment of drilling pump valve cracks can be achieved using a BW250 drilling pump vibration-acoustic test bench. This test bench may include a BW250 drilling pump, gearbox, circulating water tank, inlet and outlet pipelines, multiple accelerometers, a sound pressure sensor, a data acquisition card, and a computer. Multiple accelerometers are installed at different locations on the drilling pump cylinders or hydraulic end structures to collect the local vibration responses corresponding to multiple cylinders; the sound pressure sensor is positioned near the drilling pump to collect the global acoustic response during pump operation. The data acquisition card is used to simultaneously acquire multi-channel vibration and acoustic signals, and the computer is used to perform data storage, preprocessing, and quantitative crack assessment.

[0032] During the acquisition phase, multi-channel vibration and acoustic signals are acquired during the operation of the drilling pump to obtain operational status data. This data may include data from multiple pump cylinder vibration channels, data from one or more acoustic channels, and auxiliary information such as timestamps, sampling frequencies, operating speeds, and load conditions corresponding to the acquisition process. The multi-channel vibration signals can cover the local responses of different pump cylinder or hydraulic end structural locations, while the acoustic signals can cover the overall acoustic radiation state of the drilling pump during operation.

[0033] The pump cylinder position information corresponding to the valve to be evaluated is obtained, and the vibration channel of the corresponding pump cylinder is determined from the multi-channel vibration signal based on the pump cylinder position information to obtain the target vibration channel. The pump cylinder position information can be determined from equipment logs, sensor installation records, test calibration information, or diagnostic task configuration files. For example, valve samples with different pre-crack lengths can be installed at the valve assembly position corresponding to a specified pump cylinder. If the pre-cracked valve sample is installed at the valve assembly position corresponding to the third pump cylinder, then the vibration channel of the third pump cylinder is determined as the target vibration channel. In this way, the target vibration signal maintains a correspondence with the structural position of the valve to be evaluated.

[0034] The vibration signal corresponding to the target vibration channel is extracted from the data collected during operation to obtain the target vibration signal. The target vibration signal can be the original acceleration signal or the vibration signal after DC removal, filtering, or amplitude normalization. The acoustic signal can be the global acoustic channel signal collected by the sound pressure sensor or the acoustic signal fused from multiple acoustic sensors.

[0035] The target vibration and acoustic signals are time-aligned and amplitude preprocessed to obtain a synchronous dual-mode signal. Time alignment can be achieved based on the acquisition card's synchronization clock, trigger timestamp, or sampling point index; amplitude preprocessing can include mean removal, bandpass filtering, normalization, normalization, or outlier truncation. The target vibration and acoustic signals in the synchronous dual-mode signal correspond to the same drilling pump operation segment in the time dimension, providing a data foundation for subsequent window matching and mode fusion.

[0036] A dual-modal evaluation signal is constructed based on the synchronous dual-modal signal. This dual-modal evaluation signal can be stored in two channels, one for the target vibration signal and the other for the acoustic signal; alternatively, it can be stored in a data structure containing both vibration and acoustic components. This dual-modal evaluation signal is then input into the subsequent windowed processing flow.

[0037] In this embodiment, by using multi-channel acquisition, pump cylinder position information matching, target vibration channel extraction, and synchronous processing of vibration and acoustic signals, the dual-modal evaluation signal can simultaneously possess the local mechanical response corresponding to the valve to be evaluated and the overall acoustic response of the drilling pump, thereby improving the pertinence and stability of crack severity assessment from the data source level.

[0038] In one embodiment, the dual-modal evaluation signal is windowed to obtain a dual-modal signal window. The dual-modal evaluation signal includes target vibration signals and acoustic signals, both of which have the same sampling time reference after synchronous processing. Windowing is used to divide long-term signals during continuous operation into fixed-length samples suitable for model input.

[0039] The target vibration signal in the dual-modal evaluation signal is segmented according to a preset window length and a preset window step size to obtain a vibration window sequence. The preset window length can be determined based on the sampling frequency, drilling pump operating speed, valve impact interval, and computational resources. For example, when the sampling frequency is 25600 Hz, the preset window length can be set to 4096 sampling points. The preset window step size can be the same as the window length to achieve non-overlapping segmentation; or it can be smaller than the window length to achieve overlapping segmentation. Each vibration window covers a short operating segment, which can include valve assembly opening and closing impacts, local vibration attenuation, and structural response changes.

[0040] The acoustic signal in the dual-modal evaluation signal is segmented according to the same window position as the target vibration signal to obtain an acoustic window sequence. That is, the starting and ending sampling points and window step size of the acoustic window are consistent with the corresponding vibration window, ensuring that the vibration window and acoustic window at the same window position correspond to the same time segment during the drilling pump operation. This processing method avoids erroneous fusion caused by temporal misalignment between modes.

[0041] By pairing vibration and acoustic windows at the same window positions, a bimodal signal window is obtained. Each bimodal signal window includes one vibration component and one acoustic component. The vibration component corresponds to the local structural response of the target pump cylinder, and the acoustic component corresponds to the global acoustic response within the same running time segment. The bimodal signal window can be used as input samples for subsequent two-stream coding models.

[0042] In one optional implementation, vibration and acoustic signals are acquired at a sampling frequency of 25600 Hz, with each raw signal file recorded for 10 s. Each raw signal file can be divided into a fixed-length window of 4096 sampling points, with no overlap between windows. Thus, each window includes one vibration component and one acoustic component, forming a bimodal signal window. This window length is sufficient to cover the short-time characteristics of the dynamic response of valve impact and pump operation, while maintaining the computational efficiency of model training and inference processes.

[0043] When partitioning training, validation, and testing data for a model, data can be partitioned at the file level. Specifically, all bimodal signal windows from the same original signal file are assigned to the same subset of data in the training, validation, or testing sets, instead of randomly assigning windows from the same original signal file to different subsets. This approach reduces information leakage caused by random window-level partitioning, allowing the model performance to better reflect its generalization ability to newly acquired files and real-world engineering data.

[0044] In this embodiment, the target vibration signal and acoustic signal are segmented and paired by using the same window position, so that the vibration component and acoustic component in the dual-modal signal window have a time correspondence; the continuous signal is converted into a stable input sample by using a fixed-length window; and the risk of window-level data leakage can be reduced by using file-level data partitioning, thereby improving the reliability and engineering applicability of the crack severity assessment results.

[0045] In one embodiment, the vibration and acoustic components within a bimodal signal window are feature-encoded separately using a parameter-non-shared temporal feature coding network to obtain vibration feature word sequences and acoustic feature word sequences. The temporal feature coding network is used to convert the original time-series signal into a word sequence suitable for attention coding.

[0046] The vibration components are input into a first temporal feature encoding network for local temporal feature extraction, resulting in a vibration local feature sequence. The first temporal feature encoding network may include a convolutional input layer, multiple one-dimensional residual convolutional blocks, and an adaptive average pooling layer. The convolutional input layer is used to extract short-term impacts and local vibration patterns from the vibration waveform; the one-dimensional residual convolutional blocks are used to expand the receptive field along the time axis and alleviate gradient degradation during deep network training; the adaptive average pooling layer is used to convert intermediate features of different lengths or downsampling scales into a fixed number of vibration features.

[0047] The vibration feature sequence is subjected to feature aggregation and word mapping to obtain a vibration feature word sequence. Feature aggregation can be achieved through adaptive average pooling, temporal segmented pooling, or convolutional downsampling; word mapping can be achieved through linear mapping layers or one-dimensional convolutional mapping layers. Each vibration feature word in the vibration feature word sequence represents the vibration feature of a local time segment or a local response region.

[0048] The acoustic components are input into a second temporal feature encoding network for local temporal feature extraction, resulting in an acoustic local feature sequence. The network parameters of the first and second temporal feature encoding networks are not shared. The second temporal feature encoding network can have the same or similar network form as the first temporal feature encoding network, but its parameters are trained independently. The fluid noise, structural acoustic radiation, and propagation path coupling characteristics in the acoustic components differ from the local mechanical impact characteristics in the vibration components; therefore, using an encoding network with non-shared parameters allows the acoustic components to form feature representations suitable for their own signal characteristics.

[0049] Feature aggregation and lexical mapping are performed on the acoustic local feature sequence to obtain an acoustic feature lexical sequence. Optionally, the number of feature lexical units for each modality can be set to 64, and the embedding dimension of each feature lexical unit can be set to 128. In this case, the vibration component is encoded as 64 vibration feature lexical units, and the acoustic component is encoded as 64 acoustic feature lexical units. The two are then concatenated to form a two-stream feature lexical sequence containing 128 feature lexical units.

[0050] In this embodiment, a first temporal feature encoding network is used to extract local temporal features from the vibration component, and a second temporal feature encoding network is used to extract local temporal features from the acoustic component. This enables the vibration mode and the acoustic mode to learn feature representations suitable for their own physical sources and noise characteristics, respectively. A fixed number of feature word sequences are obtained through feature aggregation and word mapping, which can provide structured input for subsequent attention encoding and reduce the feature weakening problem caused by early mixing of heterogeneous modes.

[0051] In one embodiment, a dual-stream feature word sequence is constructed based on the vibration feature word sequence, the acoustic feature word sequence, modal identification information, and temporal location information. This dual-stream feature word sequence is used as input to the crack evaluation coding model and serves as the initial historical representation for the historical representation pool.

[0052] Vibration modal identification information is generated based on vibration modes, and acoustic modal identification information is generated based on acoustic modes. The modal identification information can be a learnable modal embedding or a modal encoding vector with the same dimension as the feature words. Vibration modal identification information is added to each vibration feature word in the vibration feature word sequence, and acoustic modal identification information is added to each acoustic feature word in the acoustic feature word sequence, enabling the crack assessment coding model to identify the physical mode to which different feature words belong.

[0053] Temporal position information is generated based on the arrangement of each feature word in the vibration and acoustic feature word sequences. This temporal position information can be a learnable position embedding, a sinusoidal position code, or other positional encoding forms. The temporal position information is used to preserve the temporal order of each feature word within the original signal window, enabling subsequent multi-head self-attention to model the dependencies between different time positions.

[0054] Vibration modal identifiers and corresponding temporal location information are added to the vibration feature lexical sequence to obtain the vibration-enhanced lexical sequence. The addition method can be word-by-word addition, linear mapping after concatenation, or other fusion methods that maintain consistent lexical dimensions. The vibration-enhanced lexical sequence simultaneously contains local vibration features, vibration modal identifiers, and temporal location information.

[0055] Acoustic modal identifiers and their corresponding temporal location information are added to the acoustic feature word sequence to obtain the acoustic enhancement word sequence. The acoustic enhancement word sequence simultaneously contains acoustic local features, acoustic modal identifiers, and temporal location information. Since the vibration modal identifiers and acoustic modal identifiers are different from each other, the subsequent crack assessment coding model can distinguish between vibration feature words and acoustic feature words within the same dual-stream feature word sequence.

[0056] The vibration-enhancing lexical sequence and the acoustic-enhancing lexical sequence are concatenated according to a preset lexical arrangement order to obtain a dual-stream feature lexical sequence. The preset lexical arrangement order can be set to arrange the vibration-enhancing lexical sequence first, followed by the acoustic-enhancing lexical sequence. A fixed arrangement order allows the subsequent modal hybrid query construction process to distinguish between vibration and acoustic modes based on lexical positions. For example, when each mode includes 64 feature lexicals, the first 64 feature lexicals in the dual-stream feature lexical sequence correspond to the vibration mode, and the last 64 feature lexicals correspond to the acoustic mode.

[0057] In this embodiment, by adding modal identification information and temporal position information to the vibration feature word sequence and the acoustic feature word sequence respectively, and splicing them according to the preset word arrangement order, the modal source information and temporal position information can be retained in the unified encoding sequence at the same time, providing a clear data structure foundation for subsequent modality-specific historical retrieval and cross-modal attention interaction.

[0058] In one embodiment, a dual-stream feature lexical sequence is input into the crack evaluation coding model, and a historical representation pool is constructed within the crack evaluation coding model. The historical representation pool is used to store historical representations generated by the crack evaluation coding model at different coding stages, enabling the current coding stage to selectively reuse historical information at different depths based on modal hybrid queries.

[0059] The dual-stream feature lexical sequence is added to the historical representation pool as the initial historical representation, resulting in the initial historical representation pool. The initial historical representation retains the original local features, modal identification information, and temporal position information of the dual-stream feature lexical sequence. For quantitative assessment of drilling pump valve cracks, the initial historical representation may contain local impact, short-time pulse, and original acoustic disturbance information. Therefore, adding it to the historical representation pool is beneficial for reusing shallow representations when needed in subsequent encoding stages.

[0060] In the current encoding stage of the crack assessment coding model, attention encoding is performed based on the current input representation to obtain a self-attention output representation. Attention encoding can be implemented using a multi-head self-attention approach to model global dependencies between vibration feature terms, between acoustic feature terms, and between vibration and acoustic feature terms. Through multi-head self-attention, the model can identify correlation patterns related to crack severity across different time locations and modes.

[0061] The self-attention output representation is added to the initial historical representation pool to obtain the first updated historical representation pool. The self-attention output representation contains intermediate features after intra-modal and inter-modal interactions. After being added to the historical representation pool, subsequent stages can retrieve the enhanced cross-modal features from it.

[0062] A feedforward transformation is performed on the self-attention output representation to obtain the feedforward output representation. This feedforward transformation can be implemented using a multilayer perceptron, which performs nonlinear mapping and channel dimension transformation on the self-attention output representation. The feedforward output representation typically contains higher-level abstract features and can be used to characterize crack severity, operating conditions, and modal coupling relationships.

[0063] The feedforward output representation is added to the first updated historical representation pool to obtain the second updated historical representation pool, which is then used as the historical representation pool for subsequent encoding stages. Thus, the historical representation pool gradually accumulates the initial representation, self-attention output representation, and feedforward output representation as the encoding stages progress. Historical representations at different encoding stages have different semantic levels, allowing subsequent modal queries to adaptively select shallow, medium, or deep representations.

[0064] In this embodiment, by continuously building and updating the historical representation pool in the crack assessment coding model, the model no longer relies on a fixed identity residual path, but can retain and reuse multi-level historical information at different coding stages. This approach is beneficial for simultaneously preserving the short-term local impact characteristics caused by valve cracks and the global fusion characteristics across operating conditions, thereby improving the stability of crack severity assessment.

[0065] In one embodiment, a vibration modal mixture query is constructed for the vibration feature word sequence, and an acoustic modal mixture query is constructed for the acoustic feature word sequence. The modal mixture query is used to retrieve historical representations from the historical representation pool that match the current modality and the current sample state.

[0066] Input-related vibration queries are obtained by extracting relevant features from the vibration feature word sequence. Input-related feature extraction can include operations such as pooling, linear mapping, normalization, and nonlinear transformation. For example, average pooling or attention pooling can be performed on the vibration feature word sequence to obtain global vibration features, which are then used to generate input-related vibration queries through a vibration query mapping network. The input-related vibration queries reflect the crack state and operating condition characteristics exhibited by the vibration modes within the current dual-modal signal window.

[0067] Learnable vibration queries corresponding to vibration modes are obtained, and the input relevant vibration queries and learnable vibration queries are weighted and combined according to vibration query gating parameters to obtain a hybrid vibration mode query. The learnable vibration query represents the global retrieval preference for vibration modes formed during model training, expressing the prior selection tendency of vibration modes towards different historical representations. The vibration query gating parameters can be obtained through learnable parameters and activation functions, and are used to adjust the contribution ratio of input relevant vibration queries and learnable vibration queries in the hybrid vibration mode query.

[0068] Input-related acoustic queries are obtained by extracting relevant features from the acoustic feature lexical sequence. These queries are generated based on the global acoustic representation of the acoustic feature lexical sequence and reflect the coupling state of acoustic radiation, fluid noise, and propagation paths within the current acoustic window. Since acoustic signals are easily affected by environmental noise, installation location, and operating conditions, input-related acoustic queries enable the historical retrieval process to remain adaptive to the current acoustic state.

[0069] Learnable acoustic queries corresponding to acoustic modes are obtained, and the input-related acoustic queries and learnable acoustic queries are weighted and combined according to acoustic query gating parameters to obtain a hybrid acoustic mode query. Specifically, the learnable vibration query, the input-related vibration query, and the vibration query gating parameters correspond to different query parameters than the learnable acoustic query, the input-related acoustic query, and the acoustic query gating parameters. In other words, vibration modes and acoustic modes do not share query parameters, enabling them to form different historical retrieval strategies.

[0070] In one embodiment, vibration mode mixed queries and acoustic mode mixed queries constitute a term-level query sequence, where vibration feature terms correspond to vibration mode mixed queries and acoustic feature terms correspond to acoustic mode mixed queries. Thus, although vibration feature terms and acoustic feature terms share the same historical representation pool, they retrieve historical representations from the pool using different query mechanisms.

[0071] In this embodiment, by weighting and combining learnable queries and input-related queries, it is possible to simultaneously utilize modal-level global retrieval preferences and the state characteristics of the current sample. By setting different query parameters and gating parameters for vibration modes and acoustic modes, it is possible to avoid different physical modes being forced to adopt the same historical retrieval strategy, thereby improving the pertinence of heterogeneous modal fusion and the reliability of crack assessment results.

[0072] like Figure 3 As shown, in one embodiment, based on vibration mode hybrid query and acoustic mode hybrid query, historical representations matching the vibration mode and historical representations matching the acoustic mode are retrieved from the historical representation pool, respectively. The retrieved historical representations are then used as residual representations in the attention encoding of the two-stream feature term sequence to obtain the fused crack representation. This process can be performed in each modality-specific hybrid query attention residual encoding block of the crack evaluation encoding model, which can be simply referred to as the MSHQ-AttnRes encoding block.

[0073] A term-level query sequence is constructed based on vibration modal hybrid queries and acoustic modal hybrid queries. The term-level query sequence corresponds one-to-one with the feature terms in the dual-stream feature term sequence, where vibration feature terms correspond to vibration modal hybrid queries, and acoustic feature terms correspond to acoustic modal hybrid queries. The term-level query sequence is used to measure the retrieval demand of each current feature term for different historical representations in the historical representation pool.

[0074] The historical representations in the historical representation pool are normalized to obtain historical key representations. Normalization can be performed using layer normalization or other normalization operations suitable for sequence representation. Since the historical representations in the pool come from different encoding stages, their amplitude distribution and semantic level may differ. Normalization can reduce the impact of differences in the scale of different historical representations on similarity calculation.

[0075] The first historical retrieval weight for each feature term is determined based on the term-level query sequence and historical key representation. The first historical retrieval weight represents the proportion of different historical representations selected by each feature term from the historical representation pool before multi-head self-attention encoding. This weight can be obtained through dot product similarity, scaling similarity, or other attention similarity calculation methods, and is normalized along the historical representation dimension.

[0076] The historical representations in the historical representation pool are weighted and aggregated according to the first historical retrieval weight to obtain the first historical residual representation. The first historical residual representation is the residual representation retrieved by the current coding block before multi-head self-attention encoding. Unlike fixed residual connections, the first historical residual representation is adaptively generated based on the current sample, the current feature terms, and the current modality, which can provide different information reuse paths for different crack states.

[0077] The first historical residual representation is used as a residual representation in the multi-head self-attention encoding of the current input representation to obtain the self-attention output representation. Multi-head self-attention encoding is used to model the dependencies between vibration feature words, between acoustic feature words, and between vibration feature words and acoustic feature words. The participation of the first historical residual representation in this process enables the attention encoding to simultaneously reference effective representations from previous historical stages when computing cross-modal interactions.

[0078] The self-attention output representation is added to the historical representation pool to obtain the updated historical representation pool. The updated historical representation pool includes both the original historical representations and the self-attention output representations generated by the current coding block, providing new candidate historical states for subsequent second historical retrieval.

[0079] The second historical representation is obtained by performing a second historical retrieval from the updated historical representation pool based on the term-level query sequence. This second historical retrieval occurs before the feedforward transform and is used to select a suitable historical representation again before the nonlinear transform. The first and second historical retrievals are located at different encoding positions; the former serves multi-head self-attention encoding, while the latter serves the feedforward transform. Therefore, they can adjust the residual paths of different sub-layers respectively.

[0080] The second historical residual representation is used as a residual representation in the feedforward transformation of the self-attention output representation to obtain the encoded output representation. The feedforward transformation can employ a multilayer perceptron to perform nonlinear processing on the self-attention output representation. The introduction of the second historical residual representation enables the feedforward transformation stage to reuse historical features that match the current modality and the current sample state.

[0081] The encoded output representation is normalized and pooled to obtain a fused crack representation. Pooling can be performed using average pooling, attention pooling, or global pooling to aggregate word-level information from the encoded output representation into a holistic crack representation. The fused crack representation includes both local impact information from the vibration modes and global sound field information from the acoustic modes, while also incorporating multi-level features of different depths from the historical representation pool.

[0082] In a preferred embodiment, the crack assessment coding model may include four MSHQ-AttnRes coding blocks, the number of attention heads may be set to 8, the multilayer perceptron expansion ratio in the feedforward transform may be set to 2.0, and the dropout may be set to 0.1. These parameters achieve a balance between model expressive power and engineering computational cost; in actual deployment, these parameters can be adjusted according to the acquisition frequency, window length, target device complexity, and computational resources.

[0083] In interpretability analysis, the first historical retrieval weight and the second historical retrieval weight can be used as the output of the modality separation attention residual weights. For example... Figure 9 As shown, the modal separation attention residual weights can characterize the selection preferences of vibrational and acoustic modes for different historical representations in different coding blocks, and can separately calculate the retrieval weights of pre-attention history retrieval and pre-feedforward history retrieval. Figure 8 As shown, the modal-level token-to-token attention matrix can display the attention relationships between vibration feature words, vibration feature words, acoustic feature words, and vibration feature words, thus helping to explain how the model performs vibration-acoustic interactions.

[0084] In this embodiment, by performing a first history retrieval before multi-head self-attention encoding and a second history retrieval before feedforward transformation, the crack assessment encoding model can establish adaptive residual paths in different sub-layers. By controlling the history retrieval process through modal hybrid query, vibration modes and acoustic modes can retrieve historical representations that match their own features. By outputting attention weights and history retrieval weights, the interpretability of the quantitative crack assessment process can also be improved.

[0085] In one embodiment, a predicted value for the severity of drilling pump valve cracks is generated based on the fused crack characterization, and a quantitative assessment result of the drilling pump valve cracks is output based on the predicted value. This process can be implemented in parallel using ordered prediction branches and regression prediction branches.

[0086] The fused crack characterization is input into the ordered prediction branch to obtain an ordered crack severity estimate. The ordered prediction branch characterizes the ordered relationship between crack levels. For example, when crack severity corresponds to 11 levels from 0 to 10 mm, the ordered prediction branch can output multiple ordered discrimination results and determine the probability that the crack severity is above the threshold of each level based on these results. Through the ordered prediction branch, the model can utilize the sequential structure of crack severity from a healthy state to different crack lengths, reducing the probability of unreasonable jumps between adjacent levels.

[0087] The fused crack characterization is input into the regression prediction branch to obtain an estimate of the severity of continuous cracks. The regression prediction branch can use a linear layer or multilayer perceptron to output continuous values ​​for directly predicting crack length or the degree of continuous damage. The regression prediction branch is suitable for outputting millimeter-level crack length estimates.

[0088] The predicted crack severity values ​​for drilling pumps and valves are obtained by fusing the estimated severity values ​​of ordered cracks and continuous cracks. The fusing process can employ a weighted fusing method, with the fusing weights determined based on the validation set error or set as preset parameters. The estimated severity values ​​of ordered cracks provide structural constraints on crack levels, while the estimated severity values ​​of continuous cracks provide fine-grained length prediction capabilities. The fusion of these two values ​​balances both ordered crack levels and quantitative continuous cracking.

[0089] During the training phase, training objectives can be set including ordered loss, regression loss, and consistency loss. Ordered loss constrains the ordered prediction branches to maintain the sequential relationship between crack severity levels; regression loss constrains the estimated severity of consecutive cracks to approximate the actual crack length; consistency loss constrains the outputs of the ordered and regression prediction branches to remain consistent. Optionally, the regression loss can use Smooth L1 loss, and the ordered loss can use an ordered loss in the form of binary cross-entropy. Model training can use the AdamW optimizer, with a batch size of 384, a training epoch count of 50, and the model parameters used during testing selected based on the mean absolute error of the validation set.

[0090] The crack length assessment and / or crack grade assessment results are determined based on the predicted crack severity values ​​for the drilling pump valve. The crack length assessment result can be a continuous millimeter value, such as 3.2 mm; the crack grade assessment result can be the closest crack grade, such as a 3 mm grade or a 4 mm grade; it can also be combined with whether the engineering tolerance output is within an acceptable range, whether the maintenance threshold has been reached, or whether the valve assembly needs to be replaced.

[0091] Based on the crack length assessment results and / or crack severity assessment results, the output of quantitative crack assessment results for drilling pumps and valves is provided. The output results may include the current crack length estimate, crack severity level, alarm information exceeding maintenance thresholds, historical trend curves, modal-level attention interpretation graphs, and historical retrieval weight interpretation graphs. Figure 6 As shown, the effectiveness of continuous prediction can be verified by the scatter distribution between the predicted value and the actual crack length; for example... Figure 7 As shown, the crack level differentiation effect can be verified by using a normalized confusion matrix; as Figure 8 As shown, the stability under different operating speeds can be verified by comparing the predicted and actual values.

[0092] In a specific experimental embodiment, the test object was a BW250 drilling pump. Radial cracks of varying lengths were pre-fabricated on the valve sample via wire cutting. A healthy state was marked as 0 mm, and crack states ranged from 1 mm to 10 mm, forming 11 ordered crack severity levels. The experiment was conducted at four operating speeds: 15 Hz, 20 Hz, 25 Hz, and 30 Hz. During data acquisition, the first three channels corresponded to the vibration responses of the three pump cylinders, and the fourth channel corresponded to the acoustic response. Since the crack was introduced into the valve assembly of the third pump cylinder, the vibration signal of the third pump cylinder and the global acoustic signal were selected as dual-modal inputs. Each signal file was divided into a 4096-point window, and the dataset was divided into training, validation, and test sets at the file level, with a ratio of 7:1:2. In the mixed-condition file-level evaluation, the training, validation, and test sets all contained samples under the four operating speed conditions.

[0093] In the above experimental embodiments, the model takes a dual-modal signal window as input and outputs predicted crack lengths within the range of 0 to 10 mm. Continuous predicted values ​​are cropped to the effective range and rounded to the nearest integer crack level when calculating classification indices; continuous predicted values ​​are directly used when calculating regression indices. Five independent runs show that the model achieves an accuracy of approximately 95.73%, with a mean absolute error of approximately 0.1520 mm, a root mean square error of approximately 0.2197 mm, and a weighted Kappa of approximately 0.9978. All test samples fall within a 1 mm engineering tolerance range. These results demonstrate that this method can achieve stable millimeter-level quantitative assessment within the complete 0 to 10 mm crack range.

[0094] In this embodiment, ordered crack severity estimates and continuous crack severity estimates are generated by ordered prediction branches and regression prediction branches, respectively, and then fused to obtain the crack severity prediction value of drilling pump valve. This can output the length of continuous cracks while maintaining the ordered structure of crack levels. By combining training loss and engineering tolerance output, the ability to distinguish between adjacent crack levels can be improved, and the needs of quantitative assessment of valve crack severity in predictive maintenance can be met.

[0095] In one embodiment, to verify the generalization ability of the quantitative assessment method for drilling pump valve cracks under unknown operating speed conditions, the four operating speeds can be denoted as the first speed condition, the second speed condition, the third speed condition, and the fourth speed condition, respectively. The first speed condition can be 15 Hz, the second speed condition can be 20 Hz, the third speed condition can be 25 Hz, and the fourth speed condition can be 30 Hz.

[0096] In the single-source migration assessment approach, a crack assessment coding model is trained using data under one rotational speed condition, and its crack severity estimation capability is tested on data under another rotational speed condition. The four rotational speeds can be paired to form 12 source-to-target domain migration tasks. This approach is used to examine the model's ability to estimate crack severity at unknown rotational speeds when only a single operating rotational speed has been encountered.

[0097] In the multi-source migration assessment approach, a crack assessment coding model is trained using data under three different rotational speed conditions, and its crack severity estimation capability is tested on data under the remaining rotational speed condition. This approach is used to examine whether multi-source operating condition coverage can mitigate the distribution shift caused by rotational speed variations. Since changes in operating speed alter valve impact intervals, hydraulic excitation frequencies, structural resonance responses, and acoustic propagation characteristics, multi-source training exposes the model to richer vibration-acoustic dynamic modes during the training phase, thereby helping to extract representations that are related to crack severity and relatively decoupled from specific rotational speeds.

[0098] In the mixed-condition file-level evaluation method, data under four different engine speeds can be divided into training, validation, and test sets at the file level, ensuring that each subset contains samples under different engine speeds. This method can evaluate the overall predictive performance of the model across multiple operating condition data coverage. Figure 7 As shown, the distribution of predicted and actual values ​​at different operating speeds can be used to demonstrate the model's crack length prediction performance under conditions of 15 Hz, 20 Hz, 25 Hz, and 30 Hz. In one example, the mean absolute error is approximately 0.077 mm at 15 Hz, approximately 0.185 mm at 20 Hz, approximately 0.102 mm at 25 Hz, and approximately 0.244 mm at 30 Hz. The accuracy within a 1 mm engineering tolerance can reach 100.0% under all speed conditions.

[0099] In this embodiment, single-source migration assessment, multi-source migration assessment, and mixed-condition file-level assessment can be used to verify the adaptability of the crack assessment coding model to changes in operating speed from different perspectives. By combining multiple historical operating condition data for training, the impact of feature distribution shift caused by speed changes on crack severity prediction can be reduced, thereby improving the robustness of quantitative crack assessment of drilling pumps and valves in actual engineering deployment.

[0100] In one embodiment, the quantitative assessment method for drilling pump valve cracks can not only output predicted values ​​of crack severity, but also output modal attention weights and historical retrieval weights to explain the contributions of vibration modes and acoustic modes in the crack assessment process.

[0101] like Figure 9As shown, the modal-level token-to-token attention matrix can be obtained by aggregating the word-level attention weights in multi-head self-attention. This attention matrix can include the attention ratio of vibration feature words to vibration feature words, the attention ratio of vibration feature words to acoustic feature words, the attention ratio of acoustic feature words to vibration feature words, and the attention ratio of acoustic feature words to acoustic feature words. Through this matrix, we can observe how the model performs vibration-acoustic interactions in different coding layers. For example, as the number of coding layers increases, the attention ratio of vibration feature words to acoustic feature words may gradually increase, indicating that acoustic information participates in the refinement of vibration representation in the deep coding stage.

[0102] like Figure 10 As shown, the modal separation attention residual weights can be obtained by statistically analyzing the first and second historical retrieval weights. The first historical retrieval weight corresponds to the historical retrieval process before multi-head self-attention encoding, and the second historical retrieval weight corresponds to the historical retrieval process before feedforward transformation. By statistically analyzing the retrieval weights of vibration feature words and acoustic feature words for different historical representations in different coding blocks, the differences in preference between the two modes for initial representations, self-attention output representations, and feedforward output representations at different depths can be analyzed.

[0103] In one implementation, vibration feature terms may be more inclined to retrieve recently enhanced structural impact representations to maintain sensitivity to local valve impacts and stiffness variations; acoustic feature terms, in certain coding layers, may be more inclined to retain initial acoustic representations or retrieve global representations at different depths to reflect the coupling information of acoustic radiation, fluid noise, and propagation paths. This difference indicates that vibration modes and acoustic modes do not share the same optimal historical retrieval trajectory, and modal hybrid queries help to form differentiated historical information reuse paths for the two modes.

[0104] In this embodiment, by outputting modal-level attention weights and modal separation attention residual weights, it is possible to explain how the model performs vibration-acoustic interaction during crack assessment and how different modes select historical representations from the historical representation pool, thereby improving the interpretability and engineering credibility of the quantitative assessment results of drilling pump valve cracks.

[0105] This embodiment provides a quantitative assessment method for drilling pump valve cracks. The method uses vibration and acoustic signals from the drilling pump operation as inputs and outputs the valve crack length or crack severity. This method is particularly suitable for crack damage assessment of hydraulic end valve assemblies in reciprocating drilling pumps, and can also be extended to damage assessment tasks of other mechanical components with localized structural impact and global acoustic radiation response.

[0106] In one embodiment, the test system includes a BW250 drilling pump, a gearbox, a circulating water tank, inlet and outlet pipelines, multiple accelerometers, a sound pressure sensor, a data acquisition card, and a computer. Multiple accelerometers are installed at different locations on the drilling pump's cylinders or hydraulic end structure to collect local vibration responses; the sound pressure sensor is positioned near the drilling pump to collect the global acoustic response during pump system operation.

[0107] Valve crack damage alters the collision, contact, leakage, and flow disturbance states between the valve body and seat, thus simultaneously affecting both vibration and acoustic signals. Vibration signals are typically more sensitive to localized impacts, changes in valve body stiffness, and pump cylinder structural responses; acoustic signals usually contain fluid noise, structural acoustic radiation, and propagation path coupling information. Although they originate from different sources, they are complementary; therefore, this embodiment simultaneously acquires vibration and acoustic signals.

[0108] Suppose the acquired dual-mode signal is: ; In the formula, This indicates the vibration signal of the target pump cylinder. This represents the acoustic signal. If the crack to be evaluated is located in the third pump cylinder valve assembly, then the vibration channel of the third pump cylinder is preferred as the acoustic signal. Acoustic signals This is the global acoustic channel for data acquisition by the sound pressure sensor.

[0109] In one optional embodiment, the vibration and acoustic signals are sampled at 25600 Hz, with each raw signal file recorded for 10 seconds. Each raw signal file is divided into a fixed-length window of 4096 sampling points, with no overlap between windows. Thus, each window contains one vibration component and one acoustic component, forming a dual-channel input. This window length is sufficient to cover the short-time characteristics of the dynamic response of valve impact and pump operation, while maintaining the computational efficiency of neural network training.

[0110] To avoid windows from the same original file appearing in both the training and test sets, data is partitioned at the file level during model training. This means all windows from the same original signal file are assigned to the same subset of the training, validation, or test sets. File-level partitioning more accurately reflects the model's generalization ability to new recorded data and reduces information leakage that might result from random window-level partitioning.

[0111] In one embodiment, the severity of the valve crack is represented by the crack length, and the label set is: In the formula, Indicates the health valve status. These represent crack lengths from 1 mm to 10 mm. For other drilling pump models or other crack calibration ranges, the label set can be expanded or scaled to the corresponding engineering calibration value.

[0112] After obtaining a fixed-length signal window, this embodiment does not simply concatenate the vibration and acoustic signals into the same feature extractor. Instead, it uses two independent one-dimensional residual network word segmenters for encoding. This allows the vibration and acoustic streams to learn local features suitable for their respective modes before deep fusion.

[0113] For any mode Its input signal The token sequence is obtained after passing through a one-dimensional residual network tokenizer: In the formula, Indicates vibration modes, Represents acoustic modes; Indicates the first One-dimensional residual network word segmenters corresponding to various modes; This indicates the number of tokens generated for each modality; Indicates the token embedding dimension. Represents a vector space.

[0114] In one specific embodiment, each tokenizer includes a convolutional input layer, multiple one-dimensional residual convolutional blocks, and an adaptive average pooling layer. The convolutional input layer is used to extract local short-term features from the original waveform; the one-dimensional residual convolutional blocks are used to expand the receptive field layer by layer along the time axis, while mitigating gradient degradation during deep network training; the adaptive average pooling layer is used to convert intermediate features of different lengths or different downsampling scales into a fixed number of tokens.

[0115] In a preferred embodiment , In other words, the vibration signal is encoded into 64 vibration tokens, and the acoustic signal is encoded into 64 acoustic tokens, resulting in a total of 128 tokens after concatenation. This setting allows for control over the computational overhead of the Transformer encoder while preserving temporal local information.

[0116] To explicitly identify the physical modality to which the token belongs, this embodiment adds learnable modality embeddings for both modalities. and Meanwhile, to maintain the positional relationships within the token sequence, learnable position embeddings are added. The initial dual-stream token sequence is: In the formula, This indicates concatenation along the token dimension; This represents the vibration token sequence after adding vibration mode embedding; This represents the acoustic token sequence after adding acoustic modal embedding; This indicates the positional embedding. The splicing order is fixed as vibration token first, then acoustic token, i.e. The fixed layout enables the subsequent modality-specific query mechanism to clearly distinguish between vibration tokens and acoustic tokens.

[0117] The residual connections in a standard Transformer typically add the current layer input directly to the sublayer output. While this fixed residual path is beneficial for gradient propagation, it cannot adaptively select historical representations of different depths based on input samples and token modes. For drilling pump and valve crack assessment, shallow features may contain localized impacts and short-duration pulses, while deep features may contain global temporal dependencies and operating condition-related information. The historical representations required for reuse are not the same for different crack lengths, rotational speeds, and modes.

[0118] Therefore, this embodiment maintains a historical representation pool in the Transformer block. Let the... The historical representation pool prior to a certain residual retrieval operation in a Transformer block is: In the formula, Indicates the first A historical state, This indicates the number of currently available historical states. Historical states can include the initial token sequence. The preceding self-attention output representation, the preceding multilayer perceptron output representation, and the preceding Transformer block output representation.

[0119] By stacking historical representations, we obtain: In the formula, Represents a history value tensor; Indicates batch size; This represents the total number of tokens after concatenating the vibration token and the acoustic token; This indicates the token dimension.

[0120] Normalizing the history value tensor yields the history key tensor: In the formula, It can be layer normalization or other normalization operations suitable for sequence representation. Normalization can reduce the impact of differences in representation amplitude at different depths on similarity calculation.

[0121] Given residual query Calculate attention weights along the historical dimension: In the formula, This represents the retrieval weight of each sample and each token in different historical states; This indicates the similarity calculation between the query and the history keys; This is a scaling factor used to improve training stability; This indicates normalization along the historical state dimension.

[0122] The residual representation after retrieval is obtained by weighting and summing the historical representations according to the historical retrieval weights: In the formula, This refers to the token representation retrieved from the historical representation pool. Indicates the first The weights of each historical state. Unlike fixed identity residuals, this representation is adaptively generated based on the current input and token type, providing differentiated information reuse paths for different crack states and modes.

[0123] Simply introducing a historical representation pool is insufficient to address the heterogeneity issue in bimodal fusion. If all tokens share the same residual query, vibration and acoustic tokens will be forced to employ the same historical retrieval strategy. However, vibration and acoustic signals have different physical meanings: vibration tokens more directly reflect local structural impact and stiffness changes, while acoustic tokens are more susceptible to fluid noise, propagation paths, and ambient sound fields. Their dependence on shallow, deep, and different sublayer historical representations also differs.

[0124] Therefore, this embodiment proposes a modality-specific hybrid query. For a vibration token, the query consists of a weighted sum of a learnable vibration query and an input-related vibration query: In the formula, This indicates a learnable vibration query, used to express the global historical retrieval preferences formed by vibration modes in the training data; This indicates an input-related query generated based on the current vibration token, used to reflect the crack state and working condition characteristics of the current sample; Represents the linear projection matrix for vibration queries; Indicates pooling operation; The vibration query gating coefficient can be obtained by applying the sigmoid function to the learnable parameters.

[0125] For acoustic tokens, the query definition is: In the formula, This indicates that acoustic queries can be learned; This indicates a query related to the input generated based on the current acoustic token; Represents the acoustic query linear projection matrix; This represents the acoustic query gating coefficient. Vibration query parameters and acoustic query parameters are not shared, thus allowing the two modes to develop different historical retrieval preferences.

[0126] Final query after token-level assembly: In the formula, the first Each query corresponds to a vibration token, and then... Each query corresponds to an acoustic token. Through this design, although vibration tokens and acoustic tokens share the same historical representation pool, they use different query mechanisms to retrieve historical representations. This preserves the deep cross-modal fusion capability while avoiding the smoothing out of modal-specific evolution patterns caused by shared queries.

[0127] In one embodiment, each MSHQ-AttnRes Transformer block includes two history retrieval operations, a multi-head self-attention layer, and a multi-layer perceptron layer. The first history retrieval occurs before the multi-head self-attention layer, and the second history retrieval occurs before the multi-layer perceptron layer.

[0128] For the Each Transformer block first performs pre-attention history retrieval: In the formula, Indicates the first Input to a Transformer block; This represents the current historical representation pool; This indicates an attention residual retrieval operation based on modality-specific hybrid queries.

[0129] Then perform multi-head self-attention: In the formula, This indicates the focus of the bulls; Representation layer normalization. Multi-head self-attention is used to model global dependencies between vibration tokens, between acoustic tokens, and between vibration tokens and acoustic tokens.

[0130] Subsequently, the multi-head self-attention output is added to the historical representation pool: Then, perform a previous multilayer perceptron historical retrieval: Finally, perform a multilayer perceptron transformation: In the formula, This represents a feedforward multilayer perceptron. Output It is also added to the historical representation pool for subsequent Transformer block retrieval. The pre-attention retrieval weights and pre-multilayer perceptron retrieval weights can be denoted as follows: and It is used to analyze the dependence of different modalities on historical representations at different depths.

[0131] In a preferred embodiment, the number of Transformer blocks The model uses 8 attention heads, an MLP scaling ratio of 2.0, and dropout of 0.1. These parameters strike a balance between model expressiveness and computational cost. In actual deployment, these parameters can be adjusted based on the acquisition frequency, window length, target device complexity, and computational resources.

[0132] After the last MSHQ-AttnRes Transformer block, the token sequence is normalized and averaged to obtain the fused representation: In the formula, This represents the token sequence output by the last Transformer block; This represents the global representation after fusion.

[0133] This embodiment uses two parallel prediction heads. The ordered prediction head outputs... One logit: In the formula, This represents the number of crack severity levels; when the crack severity level is 0 to 10 mm... ; and These are the parameters for the ordered prediction head. The crack severity estimate corresponding to the ordered prediction head is: In the formula, This represents the sigmoid function. Indicates the first An ordered logit. This form can utilize the ordered relationship between crack severity levels to reduce unreasonable jumps between adjacent levels.

[0134] The regression prediction head directly outputs the severity of continuous cracks: In the formula, and These are the parameters for the regression prediction head. The regression prediction head is helpful for obtaining estimates of crack lengths in the continuous millimeter range.

[0135] The final predicted value is obtained by fusing ordinal and regression predictions: In the formula, This represents the predicted final crack severity value; This represents the fusion coefficient. In one embodiment, In practical applications, the validation set error can be used to adjust... Adjustments will be made.

[0136] The training loss consists of three parts: In the formula, Represents the ordered loss in the form of binary cross-entropy; This represents the Smooth L1 regression loss; This represents the consistency loss, used to constrain the outputs of the ordered prediction head and the regression prediction head to be close; , and These are the loss weights.

[0137] In one specific embodiment, , , And it uses the AdamW optimizer with an initial learning rate of The weight decays to The batch size was 384, and the number of training epochs was 50. The model was selected based on the mean absolute error on the validation set, and the model parameters that performed best on the validation set were used during testing.

[0138] In one embodiment, this embodiment was conducted on a BW250 drilling pump test bench. Radial cracks of varying lengths were pre-introduced into the valve specimens via wire cutting. A healthy state was marked as 0 mm, and crack states ranged from 1 mm to 10 mm, forming 11 ordered crack severity levels. The tests were conducted at four operating speeds: 15 Hz, 20 Hz, 25 Hz, and 30 Hz. Vibration and acoustic signals for all crack levels were collected at each speed.

[0139] During data acquisition, the first three channels correspond to the vibration responses of the three pump cylinders, respectively, and the fourth channel corresponds to the acoustic response. Since the crack is introduced into the valve assembly of the third pump cylinder, the vibration signal of the third pump cylinder and the global acoustic signal are selected as dual-modal inputs. Each signal file is divided into a 4096-point window. The dataset is divided into training, validation, and test sets at the file level, with a ratio of 7:1:2.

[0140] In the mixed-condition file-level evaluation, the training, validation, and test sets all contain samples under four rotational speed conditions. The model takes a dual-flow window as input and outputs predicted crack lengths ranging from 0 to 10 mm. Continuous predicted values ​​are cropped to the effective range and rounded to the nearest integer crack class when calculating categorical indices; continuous predicted values ​​are used directly when calculating regression indices.

[0141] Experimental results show that the MSHQ-dual model in this embodiment achieves an accuracy of approximately 95.73% in five independent runs, with a mean absolute error of approximately 0.1520 mm, a root mean square error of approximately 0.2197 mm, and a weighted Kappa of approximately 0.9978. Furthermore, all test samples fall within a 1 mm engineering tolerance range. These results demonstrate that the present invention can achieve stable millimeter-level quantitative assessment within the complete 0 to 10 mm crack range.

[0142] Compared to single-vibration-flow models, dual-flow models can significantly utilize the globally complementary information in acoustic signals; compared to acoustic-only models, dual-flow models can retain the strong sensitivity of vibration signals to local valve impacts. Compared to models that do not use attention residuals or use shared query attention residuals, the modality-specific hybrid query attention residuals of this invention can better match the historical representation retrieval needs of different vibration and acoustic modes.

[0143] In one embodiment, to verify the generalization ability of the present invention under unknown operating speed conditions, four operating speeds are denoted as Hp1 to Hp4, where Hp1 is 15 Hz, Hp2 is 20 Hz, Hp3 is 25 Hz, and Hp4 is 30 Hz. Two evaluation protocols, single-source migration and multi-source migration, can be designed.

[0144] In the single-source migration protocol, the model is trained on data under one rotational speed condition and tested on data under another rotational speed condition. The four rotational speeds can be paired to form 12 source-to-target domain migration tasks. This protocol is used to examine the model's ability to estimate crack severity at unknown rotational speeds when only a single operating rotational speed has been seen.

[0145] In the multi-source migration protocol, the model is trained using data under three rotational speed conditions and tested on data under the remaining rotational speed condition. This protocol is used to examine whether multi-source operating condition coverage can alleviate rotational speed-induced domain shift. Experimental results show that multi-source migration significantly improves the target domain accuracy compared to single-source migration, indicating that this invention can learn a more stable crack-related characterization through multi-operating condition training.

[0146] From a physical perspective, changes in operating speed alter valve impact intervals, hydraulic excitation frequencies, structural resonance responses, and acoustic propagation characteristics. Multi-source training exposes the model to richer vibratory-acoustic dynamic modes during the training phase, facilitating the extraction of characterizations related to crack severity but relatively decoupled from specific operating speeds. Therefore, this invention is suitable for training with multiple historical operating condition data in engineering deployments and outputting crack severity assessment results under new operating conditions.

[0147] In one embodiment, the present invention not only outputs a predicted value of crack severity, but also outputs self-attention weights and modal separation historical retrieval weights. The self-attention weights can be aggregated into a modal-level token-to-token attention matrix, used to analyze the attention ratios of vibration tokens to vibration tokens, vibration tokens to acoustic tokens, acoustic tokens to vibration tokens, and acoustic tokens to acoustic tokens.

[0148] In one embodiment, as the number of Transformer layers increases, the proportion of attention given to the acoustic token by the vibration token gradually increases, indicating that acoustic information participates in the refinement of vibration representation at a deeper level. This phenomenon shows that the present invention does not simply superimpose two types of signals, but rather gradually forms cross-modal interaction through multi-layer self-attention.

[0149] Historical search weight and This can be used to analyze the preferences of vibration and acoustic tokens for different historical states in different Transformer blocks. In one embodiment, vibration tokens tend to retrieve recently enhanced structural impact representations, while acoustic tokens tend to retain initial acoustic representations or retrieve global representations at different depths in certain layers. This difference indicates that the two modes do not share the same optimal historical retrieval trajectory, further validating the rationality of the modality-specific hybrid query design.

[0150] Through the aforementioned interpretable outputs, maintenance personnel or model developers can observe the model's cross-modal fusion behavior and historical representation reuse behavior during crack assessment, thereby improving the model's credibility in engineering diagnostic scenarios.

[0151] In other embodiments, the number of vibration sensors can be one or more. When multiple vibration channels exist, the vibration channel corresponding to the pump cylinder where the crack is located can be selected, or multiple vibration channels can be input into multiple vibration branches separately and then fused. The acoustic sensor can also be single or multiple, and multiple acoustic sensors can form a spatial acoustic array to further improve the global sound field characterization capability.

[0152] In other embodiments, the tokenizer can employ a one-dimensional convolutional network, a lightweight residual network, a depthwise separable convolutional network, or other local encoders suitable for temporal signals. Any token that can encode vibration and acoustic signals into modality-specific token sequences can be combined with the modality-specific hybrid query attention residual mechanism of this invention.

[0153] In other embodiments, crack severity can be represented as crack length, crack grade, damage factor, remaining life percentage, or maintenance risk level. For continuous labels, regression output can be used directly; for ordered discrete labels, an ordered prediction head can be used; for engineering scenarios that require both continuous and grade values, the dual-head fusion structure of this invention can be adopted.

[0154] In other embodiments, the present invention can be deployed in a well site edge computing device, an industrial computer, or a remote server. Data acquired by the sensors can be transmitted in real time to the computing device via a data acquisition card, where preprocessing, windowing, model inference, and result output are performed. The output may include the current crack length estimate, whether the maintenance threshold has been exceeded, historical trend curves, and the corresponding attention interpretation graph.

[0155] like Figure 2 As shown, this embodiment provides a quantitative assessment system for drilling pump valve cracks. This system can be used to execute the quantitative assessment method for drilling pump valve cracks in any of the above embodiments. The system can be deployed in a computer, industrial control computer, edge computing box, server, or drilling pump condition monitoring platform. The system includes a signal acquisition module, a window construction module, a two-stream encoding module, a lexical construction module, a historical representation construction module, a modal query construction module, an attention residual encoding module, and a crack assessment output module.

[0156] The signal acquisition module acquires vibration and acoustic signals during the operation of the drilling pump. Based on the pump cylinder where the valve to be evaluated is located, it determines the target vibration signal from the vibration signals and combines the target vibration signal with the acoustic signal to form a dual-modal evaluation signal. The signal acquisition module can connect to an accelerometer, a sound pressure sensor, a data acquisition card, and an equipment operation status system. The accelerometer can be installed in different pump cylinders or hydraulic end structures, while the sound pressure sensor can be placed near the drilling pump. When the valve to be evaluated is installed in a designated pump cylinder, the signal acquisition module selects the corresponding vibration channel as the target vibration channel based on the pump cylinder location information.

[0157] The windowing module is used to window the dual-modal evaluation signal to obtain a dual-modal signal window, which includes vibration and acoustic components. The module can synchronously segment the target vibration and acoustic signals according to a preset window length and a preset window step size, and pair vibration and acoustic windows at the same window position to form a dual-modal signal window. The module can also perform preprocessing such as filtering, mean removal, amplitude normalization, and time alignment.

[0158] The dual-stream coding module is used to encode the vibration and acoustic components separately through a parameter-distributed temporal feature coding network, resulting in vibration feature word sequences and acoustic feature word sequences. The dual-stream coding module may include a first temporal feature coding network and a second temporal feature coding network. The first temporal feature coding network is used to extract local mechanical impact and structural response features from the vibration component, while the second temporal feature coding network is used to extract acoustic radiation and fluid noise features from the acoustic component.

[0159] The lexical construction module is used to construct a dual-stream feature lexical sequence based on the vibration feature lexical sequence, the acoustic feature lexical sequence, modal identifier information, and temporal position information. The lexical construction module can add vibration modal identifier information and temporal position information to the vibration feature lexical sequence, and add acoustic modal identifier information and temporal position information to the acoustic feature lexical sequence, and then concatenate them according to a preset lexical arrangement order to obtain the dual-stream feature lexical sequence.

[0160] The historical representation construction module is used to input the two-stream feature lexical sequence into the crack evaluation coding model and construct a historical representation pool within the model. The historical representation pool includes historical representations generated by the crack evaluation coding model at different coding stages, such as the initial two-stream feature lexical sequence, self-attention output representation, and feedforward output representation. The historical representation construction module can update the historical representation pool as the crack evaluation coding model runs layer by layer.

[0161] The modal query construction module is used to construct hybrid vibration modal queries for vibration feature lexical sequences and hybrid acoustic modal queries for acoustic feature lexical sequences. The hybrid vibration modal query includes learnable vibration queries and input-related vibration queries generated based on vibration feature lexical sequences, while the hybrid acoustic modal query includes learnable acoustic queries and input-related acoustic queries generated based on acoustic feature lexical sequences. The modal query construction module can also perform weighted combinations of learnable queries and input-related queries based on corresponding gating parameters.

[0162] The attention residual coding module is used to retrieve historical representations matching the vibration mode and the acoustic mode from the historical representation pool, respectively, based on vibration mode hybrid queries and acoustic mode hybrid queries. The retrieved historical representations are then used as residual representations in the attention coding of the two-stream feature term sequence to obtain the fused crack representation. The attention residual coding module may include a historical key representation generation unit, a historical retrieval weight calculation unit, a historical residual representation aggregation unit, a multi-head self-attention unit, a feedforward transformation unit, and a pooling unit. For example... Figure 2 As shown, the attention residual coding module can perform a first history retrieval before the multi-head self-attention layer and a second history retrieval before the feedforward transform layer.

[0163] The crack assessment output module generates predicted crack severity values ​​for drilling pumps and valves based on the fused crack characterization, and outputs quantitative crack assessment results for drilling pumps and valves based on these predicted crack severity values. The crack assessment output module may include an ordered prediction branch, a regression prediction branch, and a result fusion unit. The ordered prediction branch generates ordered crack severity estimates, the regression prediction branch generates continuous crack severity estimates, and the result fusion unit fuses the two and outputs crack length assessment results and / or crack grade assessment results. The crack assessment output module can also output a modal-level attention matrix and historical retrieval weights for interpretability display.

[0164] In this embodiment, the drilling pump and valve crack quantitative assessment system, through the collaborative work of the signal acquisition module, window construction module, dual-stream coding module, lexical construction module, historical characterization construction module, modal query construction module, attention residual coding module, and crack assessment output module, can realize a complete processing flow from vibration-acoustic dual-modal data acquisition to quantitative output of crack severity. By using modal-specific historical retrieval and attention residual coding, the stability of crack severity assessment under different operating conditions can be improved. By combining ordered prediction and continuous regression, crack quantitative assessment results suitable for engineering maintenance decisions can be output.

Claims

1. A method for quantitative assessment of cracks in drilling pump valves, characterized in that, The method includes: Vibration and acoustic signals during the operation of the drilling pump are acquired. The target vibration signal is determined from the vibration signals based on the pump cylinder where the valve to be evaluated is located. The target vibration signal and the acoustic signal are then combined to form a dual-mode evaluation signal. The dual-modal evaluation signal is windowed to obtain a dual-modal signal window, which includes vibration and acoustic components. The vibration component and the acoustic component are respectively encoded using a time-series feature coding network with non-shared parameters to obtain vibration feature word sequences and acoustic feature word sequences; A dual-stream feature word sequence is constructed based on the vibration feature word sequence, the acoustic feature word sequence, modal identification information, and temporal position information. The dual-stream feature word sequence is input into the crack evaluation coding model, and a historical representation pool is constructed in the crack evaluation coding model. The historical representation pool includes historical representations generated by the crack evaluation coding model at different coding stages. A vibration modal hybrid query is constructed for the vibration feature lexical sequence, and an acoustic modal hybrid query is constructed for the acoustic feature lexical sequence; wherein, the vibration modal hybrid query includes a learnable vibration query and an input-related vibration query generated based on the vibration feature lexical sequence, and the acoustic modal hybrid query includes a learnable acoustic query and an input-related acoustic query generated based on the acoustic feature lexical sequence; Based on the vibration mode hybrid query and the acoustic mode hybrid query, historical representations matching the vibration mode and historical representations matching the acoustic mode are retrieved from the historical representation pool, respectively. The retrieved historical representations are used as residual representations to participate in the attention encoding of the dual-stream feature word sequence to obtain the fused crack representation. Based on the fusion crack characterization, a predicted value for the severity of drilling pump valve cracks is generated, and a quantitative assessment result of drilling pump valve cracks is output based on the predicted value for the severity of drilling pump valve cracks.

2. The method for quantitative assessment of cracks in drilling pump valves according to claim 1, characterized in that, The process of acquiring vibration and acoustic signals during the operation of the drilling pump, determining the target vibration signal from the vibration signals based on the pump cylinder where the valve to be evaluated is located, and constructing a dual-mode evaluation signal from the target vibration signal and the acoustic signal includes: Acquire multi-channel vibration and acoustic signals during the operation of the drilling pump to obtain operational status data; Obtain the pump cylinder position information corresponding to the valve to be evaluated, and determine the vibration channel of the corresponding pump cylinder from the multi-channel vibration signal based on the pump cylinder position information to obtain the target vibration channel; The vibration signal corresponding to the target vibration channel is extracted from the data collected in the operating state to obtain the target vibration signal; The target vibration signal and the acoustic signal are time-aligned and amplitude-preprocessed to obtain a synchronous dual-mode signal; The dual-mode evaluation signal is constructed based on the synchronous dual-mode signal.

3. The method for quantitative assessment of cracks in drilling pump valves according to claim 2, characterized in that, The step of windowing the bimodal evaluation signal to obtain a bimodal signal window includes: The target vibration signal in the dual-modal evaluation signal is segmented according to the preset window length and preset window step size to obtain a vibration window sequence; The acoustic signal in the dual-modal evaluation signal is segmented according to the same window position as the target vibration signal to obtain an acoustic window sequence; The vibration window and acoustic window corresponding to the same window position are paired to obtain the dual-mode signal window.

4. The method for quantitative assessment of cracks in drilling pump valves according to claim 3, characterized in that, The vibration component and the acoustic component are respectively feature-encoded using a time-series feature coding network with non-shared parameters to obtain vibration feature word sequences and acoustic feature word sequences, including: The vibration components are input into a first temporal feature encoding network for local temporal feature extraction to obtain a vibration local feature sequence. The vibration local feature sequence is subjected to feature aggregation and word mapping to obtain the vibration feature word sequence; The acoustic components are input into a second temporal feature encoding network for local temporal feature extraction to obtain an acoustic local feature sequence; wherein, the network parameters of the first temporal feature encoding network and the second temporal feature encoding network are not shared; The acoustic local feature sequence is subjected to feature aggregation and word mapping to obtain the acoustic feature word sequence.

5. The method for quantitative assessment of cracks in drilling pump valves according to claim 4, characterized in that, The step of constructing a dual-stream feature word sequence based on the vibration feature word sequence, the acoustic feature word sequence, modal identification information, and temporal position information includes: Vibration mode identification information is generated based on the vibration mode, and acoustic mode identification information is generated based on the acoustic mode; Based on the arrangement of each feature word in the vibration feature word sequence and the acoustic feature word sequence, temporal position information is generated; The vibration mode identification information and the corresponding temporal position information are added to the vibration feature word sequence to obtain the vibration enhancement word sequence; The acoustic modality identifier information and the corresponding temporal position information are added to the acoustic feature word sequence to obtain the acoustic enhancement word sequence; The vibration-enhancing lexical sequence and the acoustic-enhancing lexical sequence are concatenated according to a preset lexical arrangement order to obtain the dual-stream feature lexical sequence.

6. The method for quantitative assessment of cracks in drilling pump valves according to claim 5, characterized in that, The step of inputting the dual-stream feature word sequence into the crack evaluation coding model and constructing a historical representation pool in the crack evaluation coding model includes: The dual-stream feature word sequence is added to the historical representation pool as the initial historical representation to obtain the initial historical representation pool; In the current encoding stage of the crack evaluation encoding model, attention encoding is performed based on the current input representation to obtain a self-attention output representation; The self-attention output representation is added to the initial historical representation pool to obtain the first updated historical representation pool; Based on the self-attention output representation, a feedforward transformation is performed to obtain the feedforward output representation; The feedforward output representation is added to the first updated history representation pool to obtain the second updated history representation pool, and the second updated history representation pool is used as the history representation pool for the subsequent encoding stage.

7. The method for quantitative assessment of cracks in drilling pump valves according to claim 6, characterized in that, The construction of a vibration modal hybrid query for the vibration feature word sequence and the construction of an acoustic modal hybrid query for the acoustic feature word sequence include: Input-related features are extracted from the vibration feature word sequence to obtain input-related vibration queries; Obtain the learnable vibration query corresponding to the vibration mode, and perform a weighted combination of the input related vibration query and the learnable vibration query according to the vibration query gating parameter to obtain the vibration mode hybrid query; Input-related features are extracted from the acoustic feature word sequence to obtain the input-related acoustic query; A learnable acoustic query corresponding to an acoustic mode is obtained, and the input-related acoustic query and the learnable acoustic query are weighted and combined according to the acoustic query gating parameters to obtain the acoustic mode hybrid query; wherein, the learnable vibration query, the input-related vibration query, and the vibration query gating parameters correspond to different query parameters as the learnable acoustic query, the input-related acoustic query, and the acoustic query gating parameters, respectively.

8. The method for quantitative assessment of cracks in drilling pump valves according to claim 7, characterized in that, The process involves retrieving historical representations matching the vibration mode and the acoustic mode from the historical representation pool, respectively, based on the vibration mode hybrid query and the acoustic mode hybrid query. These retrieved historical representations are then used as residual representations in the attention encoding of the dual-stream feature word sequence to obtain a fused crack representation, including: Construct a term-level query sequence based on the vibration mode mixed query and the acoustic mode mixed query; The historical representations in the historical representation pool are normalized to obtain historical key representations. The first historical retrieval weight corresponding to each feature word is determined based on the word-level query sequence and the historical key representation. The historical representations in the historical representation pool are weighted and aggregated according to the first historical retrieval weight to obtain the first historical residual representation. The first historical residual representation is used as a residual representation to participate in the multi-head self-attention encoding of the current input representation to obtain the self-attention output representation; The self-attention output representation is added to the historical representation pool to obtain the updated historical representation pool; Based on the term-level query sequence, a second historical retrieval is performed from the updated historical representation pool to obtain a second historical residual representation; The second historical residual representation is used as a residual representation to participate in the feedforward transformation of the self-attention output representation to obtain the encoded output representation; The encoded output representation is normalized and pooled to obtain the fusion crack representation.

9. The method for quantitative assessment of cracks in drilling pump valves according to claim 8, characterized in that, The process of generating a predicted crack severity value for the drilling pump valve based on the fusion crack characterization, and outputting a quantitative assessment result of the drilling pump valve crack based on the predicted crack severity value, includes: The fused crack characterization is input into the ordered prediction branch to obtain an estimated value of the ordered crack severity. The fused crack characterization is input into the regression prediction branch to obtain an estimate of the severity of the continuous crack. The severity estimates of ordered cracks and continuous cracks are fused together to obtain the predicted severity values ​​of the drilling pump valve cracks. The crack length assessment result and / or crack grade assessment result are determined based on the predicted crack severity value of the drilling pump valve; Based on the crack length assessment results and / or the crack grade assessment results, output the quantitative assessment results of the drilling pump valve cracks.

10. A quantitative assessment system for cracks in drilling pump valves, characterized in that, The system includes: The signal acquisition module is used to acquire vibration signals and acoustic signals during the operation of the drilling pump, determine the target vibration signal from the vibration signals according to the pump cylinder where the valve to be evaluated is located, and form a dual-mode evaluation signal by combining the target vibration signal and the acoustic signal. A window construction module is used to perform windowing processing on the dual-modal evaluation signal to obtain a dual-modal signal window, wherein the dual-modal signal window includes vibration components and acoustic components; The dual-stream coding module is used to encode the vibration component and the acoustic component separately through a time-series feature coding network with non-shared parameters, to obtain a vibration feature word sequence and an acoustic feature word sequence; The lexical construction module is used to construct a dual-stream feature lexical sequence based on the vibration feature lexical sequence, the acoustic feature lexical sequence, modal identification information, and temporal position information; The historical representation construction module is used to input the dual-stream feature word sequence into the crack evaluation coding model and construct a historical representation pool in the crack evaluation coding model. The historical representation pool includes historical representations generated by the crack evaluation coding model at different coding stages. A modal query construction module is used to construct a vibration modal hybrid query for the vibration feature lexical sequence and an acoustic modal hybrid query for the acoustic feature lexical sequence; wherein, the vibration modal hybrid query includes a learnable vibration query and an input-related vibration query generated based on the vibration feature lexical sequence, and the acoustic modal hybrid query includes a learnable acoustic query and an input-related acoustic query generated based on the acoustic feature lexical sequence; The attention residual encoding module is used to retrieve historical representations that match the vibration mode and the acoustic mode respectively from the historical representation pool based on the vibration mode mixed query and the acoustic mode mixed query, and to use the retrieved historical representations as residual representations to participate in the attention encoding of the dual-stream feature word sequence to obtain the fused crack representation; The crack assessment output module is used to generate a predicted value of the crack severity of the drilling pump valve based on the fused crack characterization, and output a quantitative assessment result of the drilling pump valve crack based on the predicted value of the crack severity of the drilling pump valve.

Citation Information

Patent Citations

  • Wind power blade early crack intelligent evaluation method and system based on acoustics and vibration signal fusion and related device

    CN121253694A

  • Valve micro-leakage detection method and device based on acoustic response analysis

    CN122360806A