Dot matrix intelligent pen with AI recording function

CN122547243APending Publication Date: 2026-08-11SHENZHEN NIUYE IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]针对现有技术的不足,本发明提供了一种带AI录音功能的点阵智能笔,解决现有的点阵智能笔,易造成笔尖摩擦噪声与人声的动态频谱混叠的问题

Benefits of technology

[0046]1. This invention uses real-time pen tip friction noise as a reference signal and synchronously acquires mixed audio signals. It combines a vibration-guided generalized sidelobe canceller to point the beam null point towards the pen tip in the spatial domain to initially suppress noise. Then, after compensating the reference signal using a secondary path model, it performs normalized least mean square adaptive filtering to eliminate linear noise residue. Based on the writing dynamic parameters and non-Gaussianity index, it selectively enables kernel adaptive filtering based on random Fourier features to fit the nonlinear transfer function. Thus, while maintaining the integrity of the human voice, it effectively separates the friction noise of dynamic spectrum aliasing, solving the problem of dynamic spectrum aliasing of pen tip friction noise and human voice in existing dot matrix smart pens.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122547243A_ABST
    Figure CN122547243A_ABST
Patent Text Reader

Abstract

The application relates to the field of intelligent writing equipment and discloses a dot matrix intelligent pen with an AI recording function, which comprises a dot matrix intelligent pen body and an AI function module arranged on the dot matrix intelligent pen body. The AI function module comprises the following units: an acquisition unit for simultaneously collecting mixed sound signals containing human voices and friction noises; a beam unit for outputting mixed signals after spatial noise reduction; a perception unit for generating dynamic parameters representing writing intensity according to writing speed and pen tip pressure; a linear unit for outputting residual error signals; a kernel filter unit for outputting pure speech estimation; and a sending unit for sending to an external terminal. By picking up the pen tip friction noise as a reference signal in real time and synchronously collecting the mixed sound signals, the vibration-guided generalized sidelobe canceller is combined to point the beam null to the pen tip direction in the spatial domain to preliminarily suppress the noise, so that the friction noise with dynamic spectral aliasing is effectively separated while the integrity of the human voice is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent writing device technology, specifically a dot matrix smart pen with AI recording function. Background Technology

[0002] A dot matrix smart pen is an intelligent writing device based on digital optical dot matrix technology. It pre-prints a dot matrix coordinate pattern that is almost invisible to the human eye on ordinary paper. A high-speed camera built into the front of the pen continuously takes pictures of the dot matrix area traversed by the pen tip at a rate of more than 100 frames per second. After image processing, data such as pen tip coordinates, pen stroke trajectory, writing speed, and pressure are extracted in real time and transmitted to an external terminal via wireless communication methods such as Bluetooth. This achieves the integration of traditional paper and pen writing with digital data acquisition. This technology has broad application prospects in education, office work, and meeting recording.

[0003] In existing dot matrix smart pens, the microphone is integrated into the pen body. When the user speaks during writing, the vibration generated by the friction between the pen tip and the paper will form strong coupled noise through solid conduction and air conduction. The main frequency of this noise is about 2kHz to 8kHz, which overlaps with the main frequency of human voice of 300Hz to 4kHz in a large area. Moreover, its spectrum changes dynamically in real time with writing speed and pen tip pressure, exhibiting non-stationary and non-linear characteristics. Conventional single-channel adaptive filtering is difficult to fit the non-linear transfer function of the noise and is easily affected by secondary path effects, resulting in unstable noise suppression effect and easy aliasing of the dynamic spectrum of pen tip friction noise and human voice. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a dot matrix smart pen with AI recording function, which solves the problem of pen tip friction noise and dynamic spectrum mixing of human voice in existing dot matrix smart pens.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a dot matrix smart pen with AI recording function, comprising a dot matrix smart pen body and an AI function module disposed on the dot matrix smart pen body, wherein the AI ​​function module comprises the following units:

[0006] Acquisition unit: used to pick up the solid-conducted vibration signal generated by the friction between the pen tip and the paper in real time as a reference signal during writing, and at the same time to collect the mixed signal containing human voice and friction noise;

[0007] Beam unit: It is used to estimate the time delay difference of friction noise reaching the acquisition point based on the cross-correlation calculation of the reference signal and the mixed signal, and dynamically adjust the blocking matrix parameters of the generalized sidelobe canceller based on the time delay difference, so that the spatial null of the beamformer points to the pen tip direction in real time, and outputs the mixed signal after spatial noise reduction.

[0008] Sensing unit: Used to calculate writing speed using the instantaneous coordinates of the pen tip and detect pen tip pressure in real time, generating dynamic parameters representing the intensity of writing based on writing speed and pen tip pressure;

[0009] Linear unit: used to obtain the compensated signal after the reference signal is convolved and compensated by the secondary path model, and to perform normalized least mean square adaptive filtering based on the compensated signal and the mixed signal to output the residual signal;

[0010] Kernel filtering unit: Based on dynamic parameters and the non-Gaussianity index of the residual signal, it uses kernel adaptive filtering based on random Fourier features to map the residual signal into a high-dimensional random feature vector and perform linear filtering to output a clean speech estimate.

[0011] Sending unit: Used to package the clean speech estimate and handwriting coordinate sequence on the same time base and send them to an external terminal.

[0012] By adopting the above technical solution, the pen tip friction noise is picked up in real time as a reference signal and the mixed signal is collected simultaneously. Combined with the vibration-guided generalized sidelobe canceller, the beam null point is pointed towards the pen tip in the spatial domain to initially suppress noise. Then, after compensating the reference signal using a secondary path model, normalized minimum mean square adaptive filtering is performed to eliminate linear noise residue. According to the writing dynamic parameters and non-Gaussianity index, kernel adaptive filtering based on random Fourier features is selectively enabled to fit the nonlinear transfer function. Thus, the friction noise of dynamic spectrum aliasing is effectively separated while maintaining the integrity of human voice. This solves the problem of dynamic spectrum aliasing of pen tip friction noise and human voice that is easy to cause in existing dot matrix smart pens.

[0013] Preferably, the acquisition unit specifically includes the following steps:

[0014] The pen tip pressure of the dot matrix smart pen is detected at a preset sampling frequency. When the pressure value is greater than the preset value and exceeds the preset time, it is in writing mode.

[0015] The solid-conducted vibration signal generated by the friction between the tip of the dot matrix smart pen and the paper is picked up at a preset sampling frequency and used as a reference signal.

[0016] A mixed audio signal containing human voice and friction noise is acquired at a preset sampling frequency using a first microphone and a second microphone arranged at intervals along the axis of the pen barrel.

[0017] Preferably, estimating the time delay difference of friction noise reaching the acquisition point includes the following steps:

[0018] The reference signal and the first signal collected by the first microphone are cross-correlated, and the time corresponding to the first peak value is extracted as the first time delay. The reference signal and the second signal collected by the second microphone are cross-correlated, and the time corresponding to the second peak value is extracted as the second time delay.

[0019] The difference between the second delay and the first delay is calculated as the relative delay difference of the friction noise arriving at the two microphones;

[0020] A blocking matrix is ​​constructed based on the relative time delay difference. The noise reference signal output by the blocking matrix retains the friction noise component in the pen tip direction and suppresses the human voice component in the user's mouth direction, thus outputting the relative time delay difference.

[0021] Preferably, the output spatially denoised mixed signal includes the following steps:

[0022] The first and second signals are input into the fixed beamformer, and the weighted sum is calculated according to the preset weight vector pointing towards the user's mouth to obtain the fixed beam output.

[0023] The first and second signals are processed by the blocking matrix to output a noise reference signal. The noise reference signal is then input into an adaptive noise canceller for normalized minimum mean square filtering to obtain a noise estimation signal.

[0024] The noise estimation signal is subtracted from the fixed beam output to obtain the spatially denoised hybrid signal.

[0025] Preferably, the sensing unit specifically includes the following steps:

[0026] The instantaneous writing speed is calculated and smoothed using the pen tip instantaneous coordinates output by the dot matrix recognition unit.

[0027] The smoothed writing speed is normalized to obtain the normalized speed, and the pen tip pressure is normalized to obtain the normalized pressure.

[0028] By weighted summing of normalization rate and normalization pressure, we obtain the dynamic parameters of writing intensity.

[0029] Preferably, obtaining the compensated signal after convolution compensation of the reference signal using a secondary path model includes the following steps:

[0030] Read the pre-stored secondary path model coefficients, and perform a convolution operation between the current sampling point and the previous preset number of sampling points of the reference signal and the secondary path model coefficients to obtain the compensation signal;

[0031] When the pen tip is off the paper and the ambient noise is below the preset decibel sound pressure level, a linear frequency modulated detection signal is emitted, and the secondary path model coefficients are updated using a recursive least squares algorithm.

[0032] Preferably, the output residual signal includes the following steps:

[0033] Select the filter order based on the dynamic parameters and determine the step size factor;

[0034] The current order of the compensation signal is used to form the input vector, which is multiplied by the current filter weight vector to obtain the linear filter output. The mixed signal is subtracted from the linear filter output to obtain the residual signal, and the filter weight vector is updated based on the residual signal and the input vector.

[0035] Preferably, the nuclear filter unit specifically includes the following steps:

[0036] The kurtosis value of the residual signal within a preset time sliding window is used as an index of non-Gaussianity.

[0037] When the kurtosis value is greater than the preset value, there is a nonlinear component. Kernel adaptive filtering based on random Fourier features is performed to output a clean speech estimate.

[0038] When the kurtosis value is less than or equal to the preset value, kernel adaptive filtering is not performed, and the residual signal is directly used as the pure speech estimate.

[0039] Preferably, the step of performing kernel adaptive filtering based on random Fourier features to output a clean speech estimate includes the following steps:

[0040] The input vector is formed by taking the number of sampling points of the current filter order of the residual signal, and then mapping the input vector into a high-dimensional random feature vector through a pre-stored random Fourier feature mapping matrix.

[0041] The kernel filter output is obtained by multiplying the high-dimensional random feature vector with the current kernel filter weight vector. The kernel filter output is then subtracted from the residual signal to obtain the clean speech estimate, and the kernel filter weight vector is updated.

[0042] Preferably, the kernel width and step size factor of the kernel adaptive filter are adaptively adjusted according to dynamic parameters, including the following steps:

[0043] The kernel width is determined based on the dynamic parameters, and the normalization step size factor is calculated based on the dynamic parameters.

[0044] Based on the determined kernel width and step size factor, high-dimensional random feature vectors are generated and kernel filter weight vectors are updated.

[0045] This invention provides a dot-matrix smart pen with AI recording function. It has the following beneficial effects:

[0046] 1. This invention uses real-time pen tip friction noise as a reference signal and synchronously acquires mixed audio signals. It combines a vibration-guided generalized sidelobe canceller to point the beam null point towards the pen tip in the spatial domain to initially suppress noise. Then, after compensating the reference signal using a secondary path model, it performs normalized least mean square adaptive filtering to eliminate linear noise residue. Based on the writing dynamic parameters and non-Gaussianity index, it selectively enables kernel adaptive filtering based on random Fourier features to fit the nonlinear transfer function. Thus, while maintaining the integrity of the human voice, it effectively separates the friction noise of dynamic spectrum aliasing, solving the problem of dynamic spectrum aliasing of pen tip friction noise and human voice in existing dot matrix smart pens.

[0047] 2. This invention activates the noise suppression pipeline only when the pen tip touches the paper through a writing state detection gating mechanism. It also dynamically adjusts the filter order, step size factor, and activation timing of kernel adaptive filtering based on the writing intensity parameters. This effectively reduces the computational load and power consumption of the main control chip while ensuring noise reduction performance, thus extending the continuous use time of the dot matrix smart pen.

[0048] 3. This invention uses an online secondary path identification mechanism to emit a detection signal in an idle state where the pen tip is off the paper and the environment is quiet. It also uses a recursive least squares algorithm to update the secondary path model in real time. This can adaptively compensate for changes in the physical transmission path caused by pen refill replacement, paper type change, or pen tip wear, thereby ensuring the convergence stability and noise reduction effect of the noise suppression system during long-term use. Attached Figure Description

[0049] Figure 1 This is an architecture diagram of the AI ​​function module of a dot matrix smart pen with AI recording function proposed in this invention.

[0050] Figure 2 This is a flowchart illustrating a method for using a dot matrix smart pen with AI recording function, as proposed in an embodiment of the present invention. Detailed Implementation

[0051] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] Example 1:

[0053] In a first embodiment of the present invention, the present invention provides a dot-matrix smart pen with AI recording function, such as... Figure 1As shown, it includes a dot matrix smart pen body and an AI function module disposed on the dot matrix smart pen body. The AI ​​function module includes the following units:

[0054] Acquisition unit: used to pick up the solid-conducted vibration signal generated by the friction between the pen tip and the paper in real time as a reference signal during writing, and at the same time to collect the mixed signal containing human voice and friction noise;

[0055] Furthermore, the acquisition unit specifically includes the following steps:

[0056] The pen tip pressure of the dot matrix smart pen is detected at a preset sampling frequency. When the pressure value is greater than the preset value and exceeds the preset time, it is in writing mode.

[0057] The solid-conducted vibration signal generated by the friction between the tip of the dot matrix smart pen and the paper is picked up at a preset sampling frequency and used as a reference signal.

[0058] A mixed audio signal containing human voice and friction noise is acquired at a preset sampling frequency using a first microphone and a second microphone arranged at intervals along the axis of the pen barrel.

[0059] Specifically, the acquisition unit is used to collect multimodal signals in real time when the dot matrix smart pen is in writing mode, providing the raw data basis for subsequent friction noise suppression and voice enhancement. The writing state is determined by the pressure sensor, and the vibration sensor and dual microphone array are activated simultaneously after confirmation to ensure strict time alignment between the reference signal and the mixing signal.

[0060] The dot-matrix smart pen has a pressure sensor installed at the end of the pen tip, which continuously detects the pressure value applied to the pen tip at a sampling frequency of 200Hz. Let the pressure sensor output value be... ,in The main control chip sets a pressure threshold for the sampling time. With duration threshold , Preferably 10 grams, The preferred time is 25 milliseconds; when continuous detection > The duration exceeded When the pen is in writing mode, it is determined that the dot matrix smart pen has entered writing mode; otherwise, it is determined to be in idle mode.

[0061] After confirming the writing state, the acquisition unit activates the vibration sensor. The vibration sensor, mounted close to the end of the pen tip, is used to pick up the solid-conducted vibration signal generated by the relative movement of the pen tip and paper. The vibration sensor is a piezoelectric thin-film sensor with a sampling frequency set to 32kHz. The output signal is recorded as follows: ,in This is a discrete-time index; the reference signal It mainly contains the solid-conductive component of pen tip friction noise.

[0062] Simultaneously, the acquisition unit activates a first microphone and a second microphone, spaced apart along the pen's axis. The first microphone is positioned near the pen tip, and the second microphone is positioned near the pen tail. Both microphones acquire sound signals at a sampling frequency of 44.1 kHz, denoted as follows: and These two signals contain a mixture of human voice, air-conducted friction noise, and solid-conducted leakage noise.

[0063] To quantify the temporal relationship between the reference signal and the mixed signal, this embodiment introduces cross-correlation calculations. The vibration sensor reference signal... With the first microphone signal The cross-correlation function is defined as: ,in This is the time delay index, corresponding to the point where the cross-correlation function reaches its maximum value. This refers to the absolute time delay of the vibration noise propagating from the pen tip to the first microphone. Similarly, calculate and cross-correlation function Absolute delay can be obtained The relative time delay difference ; The spatial orientation of the pen tip friction noise source relative to the dual microphone array is characterized and used in the construction of the blocking matrix in the subsequent beamformer.

[0064] Through the above steps, the acquisition unit completes the entire process from writing state determination to multi-channel signal acquisition, providing a time-aligned reference signal for the beam unit. First microphone signal Second microphone signal All these signals use the system clock of the main control chip as a unified time reference.

[0065] Beam unit: It is used to estimate the time delay difference of friction noise reaching the acquisition point based on the cross-correlation calculation of the reference signal and the mixed signal, and dynamically adjust the blocking matrix parameters of the generalized sidelobe canceller based on the time delay difference, so that the spatial null of the beamformer points to the pen tip direction in real time, and outputs the mixed signal after spatial noise reduction.

[0066] Furthermore, estimating the time delay difference of frictional noise reaching the acquisition point includes the following steps:

[0067] The reference signal and the first signal collected by the first microphone are cross-correlated, and the time corresponding to the first peak value is extracted as the first time delay. The reference signal and the second signal collected by the second microphone are cross-correlated, and the time corresponding to the second peak value is extracted as the second time delay.

[0068] The difference between the second delay and the first delay is calculated as the relative delay difference of the friction noise arriving at the two microphones;

[0069] A blocking matrix is ​​constructed based on the relative time delay difference. The noise reference signal output by the blocking matrix retains the friction noise component in the pen tip direction and suppresses the human voice component in the user's mouth direction, thus outputting the relative time delay difference.

[0070] Furthermore, the output spatially denoised mixed signal includes the following steps:

[0071] The first and second signals are input into the fixed beamformer, and the weighted sum is calculated according to the preset weight vector pointing towards the user's mouth to obtain the fixed beam output.

[0072] The first and second signals are processed by the blocking matrix to output a noise reference signal. The noise reference signal is then input into an adaptive noise canceller for normalized minimum mean square filtering to obtain a noise estimation signal.

[0073] The noise estimation signal is subtracted from the fixed beam output to obtain the spatially denoised hybrid signal.

[0074] Specifically, the beamforming unit is used to initially suppress friction noise in the pen tip direction in the spatial domain based on the time relationship between the reference signal and the mixed signal, thereby improving the signal-to-noise ratio of the mixed signal. The beamforming unit first estimates the time delay difference of friction noise reaching the dual microphones through cross-correlation calculation, and then dynamically adjusts the blocking matrix parameters of the generalized sidelobe canceller based on the time delay difference, so that the spatial zero point of the beamformer points to the pen tip direction in real time, and finally outputs the spatially denoised mixed signal.

[0075] To estimate the time difference of friction noise propagating from the pen tip to the two microphones, the beamforming unit calculated... and and The cross-correlation function, the cross-correlation operation is performed only within a finite time delay, for example ,in The number of sampling points corresponds to 5 milliseconds, in order to reduce the amount of computation.

[0076] Reference signal With the first microphone signal The cross-correlation function is defined as: ,in This is the time delay index; the index corresponding to the maximum value of the cross-correlation function. This refers to the absolute time delay of the vibration noise propagating from the pen tip to the first microphone. Similarly, calculate and cross-correlation function Absolute delay can be obtained The relative time delay difference In this plan, The sign and magnitude of the pen tip friction noise source reflect its orientation relative to the axis of the dual-microphone array; for example, when When positive, it indicates that the noise source is closer to the first microphone; when... When the value is negative, it indicates that the noise source is closer to the second microphone. This relative time delay difference serves as the basis for constructing the subsequent blocking matrix and is output by the beam unit to the generalized sidelobe canceller.

[0077] Next, the beam unit outputs a spatially denoised mixed signal. The generalized sidelobe canceller consists of three parts: a fixed beamformer, a blocking matrix, and an adaptive noise canceller.

[0078] The fixed beamformer receives the first microphone signal. Second microphone signal The weighted sum is calculated according to a preset weight vector pointing towards the user's mouth; let the weight vector be... ,in and As a preset constant, its value is chosen to align the main lobe of the fixed beamformer with the user's mouth direction. Therefore, the fixed beam output is: In this embodiment, and All values ​​are 0.5, meaning they are simply added together.

[0079] The blocking matrix is ​​used to filter out the desired signal components from the user's mouth direction while retaining the noise components from the pen tip direction; the construction of the blocking matrix B depends on the aforementioned relative time delay difference. The blocking matrix is ​​a 2×1 column vector. Its elements satisfy + =0, and according to The phase difference between the two signals is adjusted so that the signals incident from the user's port direction cancel each other out after processing by the blocking matrix. The output of the first and second signals after processing by the blocking matrix is: In this embodiment, ,but This output serves as a noise reference signal. .

[0080] Adaptive noise canceller for noise reference signal Adaptive filtering is performed to estimate the frictional noise component in the mixed signal; let the filter weight vector of the adaptive noise canceller be... ,in The filter order is used; in this embodiment, If the optimal value is 32, then the noise estimation signal is: The adaptive noise canceller updates the weight vector using the normalized least mean square algorithm, and its update formula is: ,in The noise reference signal vector, For error signals, Step size factor To prevent division by zero constant; in this embodiment, The preferred value is 0.05. The preferred value is 0.001; error signal It is obtained by subtracting the noise estimation signal from the fixed beam output, i.e. This error signal is the mixed signal after spatial noise reduction, denoted as... .

[0081] Through the above steps, the beamforming unit uses time delay difference estimation guided by the vibration sensor to dynamically adjust the generalized sidelobe canceller, ensuring that the spatial zero point points in real time toward the pen tip, thereby effectively suppressing friction noise in the spatial domain. The final output... The desired signal for subsequent linear units is used for further speech enhancement processing.

[0082] Sensing unit: Used to calculate writing speed using the instantaneous coordinates of the pen tip and detect pen tip pressure in real time, generating dynamic parameters representing the intensity of writing based on writing speed and pen tip pressure;

[0083] Furthermore, the sensing unit specifically includes the following steps:

[0084] The instantaneous writing speed is calculated and smoothed using the pen tip instantaneous coordinates output by the dot matrix recognition unit.

[0085] The smoothed writing speed is normalized to obtain the normalized speed, and the pen tip pressure is normalized to obtain the normalized pressure.

[0086] By weighted summing of normalization rate and normalization pressure, we obtain the dynamic parameters of writing intensity.

[0087] Specifically, the sensing unit is used to extract dynamic features of the writing process from the dot matrix recognition unit and the pressure sensor, and generate dynamic parameters that characterize the intensity of writing. These dynamic parameters are then used to adjust the order, step size, and activation threshold of the adaptive filter and the kernel adaptive filter, so as to enable the noise suppression system to follow the changes in writing behavior in real time. The sensing unit first uses the instantaneous coordinates of the pen tip output by the dot matrix recognition unit to calculate the writing speed and smooth it. Then, it normalizes the writing speed and pen tip pressure respectively, and finally obtains the dynamic parameters of writing intensity by weighted summation.

[0088] The dot matrix recognition unit outputs the instantaneous coordinates of the pen tip on the dot matrix paper at a rate of 100 to 200 frames per second, denoted as... and ,in For frame index, the time interval between two adjacent frames is denoted as . Let the first The coordinates of the frame are and , No. The coordinates of the frame are and Then instantaneous writing speed Defined as the Euclidean distance between the coordinates of two adjacent frames divided by the inter-frame time interval, the calculation formula is: In this embodiment, The frame rate is determined by the dot matrix camera; for example, when the frame rate is 100 frames per second, It takes about 10 milliseconds.

[0089] Because speed calculations between adjacent frames may contain jitter, instantaneous writing speed needs to be smoothed. A five-point moving average is used to smooth the speed. Smoothing is performed to achieve a smooth writing speed. The calculation formula is as follows: ,in ≥4, for the starting frame, fewer points can be used or the original value can be kept; the smoothed writing speed can more stably reflect the overall trend of the writing action.

[0090] To unify writing speed and pen tip pressure on the same scale, the sensing unit normalizes both smooth writing speed and pen tip pressure separately. The pen tip pressure value detected in real-time by the pressure sensor is set to... In this embodiment, a maximum writing speed is preset. The preferred setting is 20 centimeters per second, with the maximum pen tip pressure preset. The preferred pressure is 300 grams, with a preset minimum pen tip pressure. The preferred value is 10 grams, then the normalized speed and normalization pressure They are defined as follows:

[0091] ;

[0092] ;

[0093] And on The normalization process limits the writing speed and pen pressure to a dimensionless value, making subsequent weighted fusion easier.

[0094] Dynamic parameters of writing intensity The formula is obtained by weighted summation of normalized velocity and normalized pressure, and is as follows: ,in This is the speed weighting coefficient, with a value ranging from 0 to 1. In this embodiment, The preferred value is 0.6. This weighting is based on experimental observations: changes in writing speed have a more significant impact on the friction noise spectrum, while pen tip pressure mainly affects the total energy of the noise.

[0095] Through the above steps, the sensing unit converts the dot matrix coordinates and the raw data output by the pressure sensor into a dynamic parameter between 0 and 1. When the user writes slowly, and All are relatively small. Approaching 0; when the user writes quickly and forcefully, Approaching 1; this dynamic parameter serves as the basis for adjusting subsequent linear units and kernel filtering units, enabling the noise suppression system to adaptively follow changes in writing behavior in real time.

[0096] Linear unit: used to obtain the compensated signal after the reference signal is convolved and compensated by the secondary path model, and to perform normalized least mean square adaptive filtering based on the compensated signal and the mixed signal to output the residual signal;

[0097] Furthermore, the reference signal is compensated by convolution of the secondary path model to obtain the compensated signal, including the following steps:

[0098] Read the pre-stored secondary path model coefficients, and perform a convolution operation between the current sampling point and the previous preset number of sampling points of the reference signal and the secondary path model coefficients to obtain the compensation signal;

[0099] When the pen tip is off the paper and the ambient noise is below the preset decibel sound pressure level, a linear frequency modulated detection signal is emitted, and the secondary path model coefficients are updated using a recursive least squares algorithm.

[0100] Furthermore, the residual signal is output, including the following steps:

[0101] Select the filter order based on the dynamic parameters and determine the step size factor;

[0102] The current order of the compensation signal is used to form the input vector, which is multiplied by the current filter weight vector to obtain the linear filter output. The mixed signal is subtracted from the linear filter output to obtain the residual signal, and the filter weight vector is updated based on the residual signal and the input vector.

[0103] Specifically, the linear unit is used to perform first-stage adaptive filtering on the reference signal picked up by the vibration sensor after compensation by the secondary path model, and the spatial noise-reduced mixed signal output by the beamforming module to eliminate linear predictable components in the noise. The linear unit first reads the pre-stored secondary path model coefficients, and performs convolution operation on the reference signal and the model coefficients to obtain the compensation signal; then, it selects the filter order and step size factor according to the dynamic parameters output by the sensing unit, performs normalized least mean square adaptive filtering, and outputs the residual signal; in addition, the linear unit uses the built-in exciter to emit a probe signal in the idle state, and uses the recursive least squares algorithm to update the secondary path model online to adapt to changes in the physical path caused by pen refill replacement or paper type change.

[0104] Vibration sensor reference signal Mixed signal with microphone There exists a secondary path effect, which is the physical transmission path from the vibration sensor pickup point to the microphone acoustic cavity entrance; this path includes two stages: solid conduction and air propagation, and its transfer function is denoted as... Without compensation, the convergence direction of the adaptive filter will deviate from the optimal solution. Therefore, a secondary path model is introduced into the linear unit. The reference signal is pre-filtered.

[0105] The secondary path model coefficients were obtained through offline calibration. The calibration process was as follows: In a quiet room environment, the paper was run at a standard writing speed and pressure on a dot matrix paper surface, and vibration sensor signals were collected. Friction noise components in microphone signals The least squares method is used to identify lengths of... Finite impulse response filter coefficients In this embodiment, The preferred value is 64; this set of coefficients is pre-stored in the non-volatile memory of the dot matrix smart pen.

[0106] Let the reference signal be The current sampling point and the previous The vector consists of 1 sampling point. Then the compensation signal Obtained through convolution operations: The compensation signal It serves as a reference input for subsequent adaptive filters and is used for coefficient update direction correction.

[0107] To adapt to changes in the secondary path during use, the linear unit performs online model updates in an idle state where the pen tip is off the paper and the ambient noise is below a preset sound pressure level. In this scheme, the ambient noise threshold is preferably 40 dB sound pressure level. When the condition is met, the linear unit emits a linear frequency modulated probe signal through a built-in piezoelectric exciter. The frequency range of this signal covers 1 kHz to 10 kHz, and the duration is 0.2 seconds. Simultaneously, vibration sensor signal xprobe[n] and microphone signal dprobe[n] are acquired, and the secondary path model coefficients are updated using a recursive least squares algorithm. The update formula for the recursive least squares algorithm is:

[0108] ;

[0109] ;

[0110] ;

[0111] in For the gain vector, It is the inverse of the covariance matrix. Forgetting factor, It is the identity matrix; in this embodiment, The preferred value is 0.998; the updated model coefficients replace the original coefficients and are used for subsequent convolution compensation.

[0112] Next, the linear unit performs normalized least mean square adaptive filtering to output the residual signal. Let the dynamic parameter of the writing intensity output by the sensing unit be... Linear unit according to Select the filter order and step size factor ,when When <0.3, , When 0.3≤ When ≤0.7, , ;when When >0.7, , The step size factor can also be uniformly expressed as And set upper and lower limits.

[0113] Compensation signal The current The input vector consists of 1 sampling point. Let the current linear filter weight vector be... Then the linear filter output is: The spatially denoised hybrid signal output by the beamforming module is denoted as , the residual signal Defined as: The linear filter weight vector is updated using the normalized least mean square algorithm, and the update formula is as follows: ,in The energy of the input vector. To prevent division by zero constant; in this embodiment, The preferred value is 0.001; this update rule ensures that the filter step size automatically decreases when the input signal energy is large, thereby maintaining convergence stability.

[0114] Through the above steps, the linear unit completes the entire process from reference signal compensation to adaptive filtering, outputting the residual signal. The linear predictable noise in the residual signal has been significantly reduced, and the remaining components are mainly nonlinear noise residues, which are then processed by the subsequent kernel filtering module.

[0115] Kernel filtering unit: Based on dynamic parameters and the non-Gaussianity index of the residual signal, it uses kernel adaptive filtering based on random Fourier features to map the residual signal into a high-dimensional random feature vector and perform linear filtering to output a clean speech estimate.

[0116] Furthermore, the nuclear filtering unit specifically includes the following steps:

[0117] The kurtosis value of the residual signal within a preset time sliding window is used as an index of non-Gaussianity.

[0118] When the kurtosis value is greater than the preset value, there is a nonlinear component. Kernel adaptive filtering based on random Fourier features is performed to output a clean speech estimate.

[0119] When the kurtosis value is less than or equal to the preset value, kernel adaptive filtering is not performed, and the residual signal is directly used as the pure speech estimate.

[0120] Furthermore, kernel adaptive filtering based on random Fourier features is performed to output a clean speech estimate, including the following steps:

[0121] The input vector is formed by taking the number of sampling points of the current filter order of the residual signal, and then mapping the input vector into a high-dimensional random feature vector through a pre-stored random Fourier feature mapping matrix.

[0122] The kernel filter output is obtained by multiplying the high-dimensional random feature vector with the current kernel filter weight vector. The kernel filter output is then subtracted from the residual signal to obtain the clean speech estimate, and the kernel filter weight vector is updated.

[0123] Furthermore, the kernel width and step size factor of the kernel adaptive filter are adaptively adjusted according to dynamic parameters, including the following steps:

[0124] The kernel width is determined based on the dynamic parameters, and the normalization step size factor is calculated based on the dynamic parameters.

[0125] Based on the determined kernel width and step size factor, high-dimensional random feature vectors are generated and kernel filter weight vectors are updated.

[0126] Specifically, the kernel filtering unit selectively enables kernel adaptive filtering based on stochastic Fourier features according to the degree of nonlinearity of the residual signal output by the linear unit, in order to eliminate nonlinear noise components in the residual and finally output a clean speech estimate. The kernel filtering unit first calculates the kurtosis value of the residual signal within the sliding window as a non-Gaussianity index to determine whether there are significant nonlinear components. When nonlinear components exist, the kernel width and step size factor are adaptively adjusted according to the dynamic parameters output by the perception unit, and the residual signal is mapped into a high-dimensional random feature vector using a pre-stored stochastic Fourier feature mapping matrix, and linear filtering is performed in the high-dimensional feature space. When the nonlinear components are not significant, kernel adaptive filtering is directly bypassed to save computational resources.

[0127] The residual signal output by the linear unit is denoted as ,in This is a discrete-time index; the residual signal may still contain some nonlinear noise components, which typically exhibit non-Gaussian distribution characteristics. Kurtosis is a commonly used statistic to measure the non-Gaussianity of a signal, defined as the ratio of the fourth-order cumulant to the square of the second-order cumulant. In this embodiment, the kernel filter unit calculates... Kurtosis value within the sliding window The calculation formula is: ,in The length of the sliding window. The mean of the residual signal within the window is used in this embodiment. The preferred number of sampling points corresponds to 200 milliseconds, i.e., when the sampling rate is 44.1 kHz. Approximately 8820, kurtosis value This reflects the tail thickness of the signal distribution: for a Gaussian distribution, The theoretical value is 3; when When the value is greater than 3, the signal exhibits a super-Gaussian distribution, indicating the presence of spike-pulse nonlinear components; in this embodiment, a preset kurtosis threshold is used. Preferably 3.5; when When the value is greater than 3.5, it is determined that there is a significant nonlinear component in the residual signal, and the kernel filtering unit enables kernel adaptive filtering based on random Fourier features; when When the value is ≤3.5, the nonlinear component can be ignored, and the kernel filter unit directly sets the output. This skips subsequent filtering steps.

[0128] When a nonlinear component is determined to exist, the kernel filtering unit performs kernel adaptive filtering based on stochastic Fourier features. This process includes input vector construction, stochastic Fourier feature mapping, linear filtering and weight vector update, and adaptive adjustment of kernel width and step size factors.

[0129] First, the kernel filtering unit extracts the residual signal. The current The input vector consists of 1 sampling point. .in As the input dimension of the kernel filter, based on the dynamic parameters output by the sensing unit. Determined; in this embodiment, when <0.3 =32, when 0.3≤ ≤0.7 =64, when >0.7 =128; This dynamic adjustment strategy allows the filter to have sufficient degrees of freedom to fit complex nonlinear mappings when writing drastic, nonlinear enhancements.

[0130] Next, the kernel filtering unit uses a pre-stored random Fourier feature mapping matrix to map the input vector to a high-dimensional feature space. Random Fourier features are a technique for approximating kernel functions; their basic principle is to construct explicit feature mappings through random sampling, making the linear inner product in the high-dimensional feature space approximate the translation-invariant kernel function of the original input space. In this embodiment, a Gaussian kernel function is selected. ,in Given the kernel width, the random Fourier feature map is constructed as follows: Let the mapping dimension be D, first generate a random weight matrix. and random bias vector ,in Each element follows a mean of 0 and a variance of . Gaussian distribution, Each element follows a range If the input vector is uniformly distributed on the input vector, then... The mapped high-dimensional random feature vector is: ,in It is an element-wise cosine function; in this embodiment, The preferred value is 512; this mapping satisfies That is, the inner product in the characteristic space approximates the Gaussian kernel function.

[0131] The kernel filter unit performs linear filtering in a high-dimensional feature space. Let the weight vector of the kernel filter be... The initial value is a zero vector; kernel filter output. The dot product of the weight vector and the eigenvector: The pure speech estimate e[n] is obtained by subtracting the kernel filter output from the residual signal: The weight vector is updated using the normalized least mean square rule, and the update formula is: ,in Step size factor To prevent the division into zero constants, a value of 0.001 is preferred.

[0132] nuclear width and step size factor All based on dynamic parameters Adaptive adjustment. Core width. The adjustment strategy is as follows: as the intensity of writing increases, the nonlinearity of friction noise increases, requiring a smaller kernel width to improve local resolution. In this embodiment, the kernel width is determined by the following formula: ,in The initial kernel width, In this embodiment, the attenuation coefficient is used. The preferred value is 4.5. The preferred value is 1, the step size factor. Similarly, with dynamic parameters Negative correlation is used to ensure the stability of the filter during high-speed writing. ,in To preset the maximum step size, The attenuation coefficient is used in this scheme. The preferred value is 0.5. The preferred value is 2.5; it should be noted that in the denominator... This refers to the denominator term in the aforementioned update formula; its repetition here is solely to emphasize the normalization property of the step size factor. In practical implementation, the step size factor can be directly adopted. It is also supplemented with amplitude limiting to avoid problems such as division by zero and insufficient energy.

[0133] Through the above steps, when the nonlinear components are significant, the kernel filtering unit uses random Fourier features to map the input to a high-dimensional space and performs kernel adaptive filtering, effectively eliminating the nonlinear noise residue in the residuals. When the nonlinear components are not significant, direct bypass filtering is used to reduce computational overhead, resulting in a clean speech estimate as the final output. It is sent to the sending module for packaging and transmission.

[0134] Sending unit: Used to package the clean speech estimate and handwriting coordinate sequence on the same time base and send them to an external terminal.

[0135] Specifically, the sending unit is used to package the pure speech estimation and handwriting coordinate sequence with the same time base and send them to the external terminal. The sending unit obtains the system clock of the main control chip as a unified time base and adds a timestamp to each frame of speech data and each group of handwriting data. The pure speech estimation is encoded in 16kHz sampling rate and 16-bit PCM format, and is packaged into a frame every 20 milliseconds. The handwriting coordinate sequence is collected at a sampling rate of 200Hz, with a data point every 5 milliseconds, including the horizontal coordinate, vertical coordinate, pressure and speed. The sending unit transmits the speech frame and handwriting data packet synchronously to the external terminal through the Bluetooth isochronous channel, so that the terminal can achieve precise synchronization of speech playback and handwriting playback based on the timestamp.

[0136] Example 2:

[0137] In a second embodiment of the present invention, the present invention provides a method for using a dot-matrix smart pen with AI recording function, such as... Figure 2 As shown, it includes the following steps:

[0138] In writing mode, the solid-conducted vibration signal generated by the friction between the pen tip and the paper is picked up in real time as a reference signal, and a mixed signal containing human voice and friction noise is also collected.

[0139] Based on the cross-correlation calculation of the reference signal and the mixed signal, the time delay difference of the friction noise reaching the acquisition point is estimated, and the blocking matrix parameters of the generalized sidelobe canceller are dynamically adjusted based on the time delay difference, so that the spatial null point of the beamformer points to the pen tip direction in real time, and the mixed signal after spatial noise reduction is output.

[0140] The writing speed is calculated using the instantaneous coordinates of the pen tip, and the pen tip pressure is detected in real time. Based on the writing speed and pen tip pressure, dynamic parameters representing the intensity of writing are generated.

[0141] The reference signal is convolved and compensated by a secondary path model to obtain the compensated signal. Then, normalized least mean square adaptive filtering is performed on the compensated signal and the mixed signal to output the residual signal.

[0142] Based on the dynamic parameters and the non-Gaussianity index of the residual signal, kernel adaptive filtering based on random Fourier features is used to map the residual signal into a high-dimensional random feature vector and perform linear filtering to output a clean speech estimate.

[0143] The pure speech estimation and handwriting coordinate sequence are packaged together on the same time base and sent to an external terminal.

[0144] In a classroom lesson, the teacher used a dot-matrix smart pen to explain math examples on a dot-matrix notebook, verbally explaining the solution steps while writing. However, due to the noise generated by the friction between the pen tip and the paper mixing with the teacher's voice, traditional recording equipment struggled to clearly capture the teacher's speech, resulting in students not being able to hear the explanation clearly when playing it back after class. To solve this problem, this invention provides a method for using a dot-matrix smart pen with AI recording functionality, the process of which is as follows: Figure 2 As shown. The specific implementation process of this method is as follows:

[0145] First, in writing mode, the dot matrix smart pen uses a vibration sensor at the end of the pen tip to pick up solid-conducted friction noise in real time as a reference signal, and at the same time uses dual microphones to collect a mixed signal containing human voice and friction noise.

[0146] Then, the time delay difference is estimated based on the cross-correlation operation between the reference signal and the mixed signal, and the blocking matrix parameters of the generalized sidelobe canceller are dynamically adjusted so that the spatial null of the beamformer points to the pen tip direction, and the spatially denoised mixed signal is output.

[0147] Next, the writing speed is calculated using the instantaneous coordinates of the pen tip output by the dot matrix recognition unit, and dynamic parameters characterizing the intensity of writing are generated by combining the pen tip pressure detected by the pressure sensor.

[0148] Then, the reference signal is compensated by convolution of the secondary path model and then normalized least mean square adaptive filtering is performed to output the residual signal.

[0149] Then, based on the dynamic parameters and the non-Gaussianity index of the residual signal, kernel adaptive filtering based on random Fourier features is selectively enabled to eliminate nonlinear noise residues and output clean speech estimation.

[0150] Finally, the pure speech estimation and handwriting coordinate sequence are packaged on the same time base and sent to the student terminal via Bluetooth.

[0151] Using the above method, the teacher's voice during the writing process is clearly captured, and students can hear clear explanations and see handwriting animations simultaneously when playing back the recording, effectively improving the after-class review effect.

[0152] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A dot matrix smart pen with AI recording function, comprising a dot matrix smart pen body and an AI function module arranged on the dot matrix smart pen body, characterized in that: The AI ​​functional module includes the following units: Acquisition unit: used to pick up the solid-conducted vibration signal generated by the friction between the pen tip and the paper in real time as a reference signal during writing, and at the same time to collect the mixed signal containing human voice and friction noise; Beam unit: It is used to estimate the time delay difference of friction noise reaching the acquisition point based on the cross-correlation calculation of the reference signal and the mixed signal, and dynamically adjust the blocking matrix parameters of the generalized sidelobe canceller based on the time delay difference, so that the spatial null of the beamformer points to the pen tip direction in real time, and outputs the mixed signal after spatial noise reduction. Sensing unit: Used to calculate writing speed using the instantaneous coordinates of the pen tip and detect pen tip pressure in real time, generating dynamic parameters representing the intensity of writing based on writing speed and pen tip pressure; Linear unit: used to obtain the compensated signal after the reference signal is convolved and compensated by the secondary path model, and to perform normalized least mean square adaptive filtering based on the compensated signal and the mixed signal to output the residual signal; Kernel filtering unit: Based on dynamic parameters and the non-Gaussianity index of the residual signal, it uses kernel adaptive filtering based on random Fourier features to map the residual signal into a high-dimensional random feature vector and perform linear filtering to output a clean speech estimate. Sending unit: Used to package the clean speech estimate and handwriting coordinate sequence on the same time base and send them to an external terminal. 2.The dot matrix smart pen with AI recording function of claim 1, wherein: The acquisition unit specifically includes the following steps: The pen tip pressure of the dot matrix smart pen is detected at a preset sampling frequency. When the pressure value is greater than the preset value and exceeds the preset time, it is in writing mode. The solid-conducted vibration signal generated by the friction between the tip of the dot matrix smart pen and the paper is picked up at a preset sampling frequency and used as a reference signal. A mixed audio signal containing human voice and friction noise is acquired at a preset sampling frequency using a first microphone and a second microphone arranged at intervals along the axis of the pen barrel.

3. A dot-matrix smart pen with AI recording function according to claim 1, characterized in that: The estimation of the time delay difference of friction noise reaching the acquisition point includes the following steps: The reference signal and the first signal collected by the first microphone are cross-correlated, and the time corresponding to the first peak value is extracted as the first time delay. The reference signal and the second signal collected by the second microphone are cross-correlated, and the time corresponding to the second peak value is extracted as the second time delay. The difference between the second delay and the first delay is calculated as the relative delay difference of the friction noise arriving at the two microphones; A blocking matrix is ​​constructed based on the relative time delay difference. The noise reference signal output by the blocking matrix retains the friction noise component in the pen tip direction and suppresses the human voice component in the user's mouth direction, thus outputting the relative time delay difference.

4. A dot-matrix smart pen with AI recording function according to claim 1, characterized in that: The output spatially denoised mixed signal includes the following steps: The first and second signals are input into the fixed beamformer, and the weighted sum is calculated according to the preset weight vector pointing towards the user's mouth to obtain the fixed beam output. The first and second signals are processed by the blocking matrix to output a noise reference signal. The noise reference signal is then input into an adaptive noise canceller for normalized minimum mean square filtering to obtain a noise estimation signal. The noise estimation signal is subtracted from the fixed beam output to obtain the spatially denoised hybrid signal.

5. A dot-matrix smart pen with AI recording function according to claim 1, characterized in that: The sensing unit specifically includes the following steps: The instantaneous writing speed is calculated and smoothed using the pen tip instantaneous coordinates output by the dot matrix recognition unit. The smoothed writing speed is normalized to obtain the normalized speed, and the pen tip pressure is normalized to obtain the normalized pressure. By weighted summing of normalization rate and normalization pressure, we obtain the dynamic parameters of writing intensity.

6. A dot-matrix smart pen with AI recording function according to claim 1, characterized in that: The process of obtaining the compensated signal by convolution compensation of the reference signal using a secondary path model includes the following steps: Read the pre-stored secondary path model coefficients, and perform a convolution operation between the current sampling point and the previous preset number of sampling points of the reference signal and the secondary path model coefficients to obtain the compensation signal; When the pen tip is off the paper and the ambient noise is below the preset decibel sound pressure level, a linear frequency modulated detection signal is emitted, and the secondary path model coefficients are updated using a recursive least squares algorithm.

7. A dot-matrix smart pen with AI recording function according to claim 1, characterized in that: The output residual signal includes the following steps: Select the filter order based on the dynamic parameters and determine the step size factor; The current order of the compensation signal is used to form the input vector, which is multiplied by the current filter weight vector to obtain the linear filter output. The mixed signal is subtracted from the linear filter output to obtain the residual signal, and the filter weight vector is updated based on the residual signal and the input vector.

8. A dot-matrix smart pen with AI recording function according to claim 1, characterized in that: The nuclear filtering unit specifically includes the following steps: The kurtosis value of the residual signal within a preset time sliding window is used as an index of non-Gaussianity. When the kurtosis value is greater than the preset value, there is a nonlinear component. Kernel adaptive filtering based on random Fourier features is performed to output a clean speech estimate. When the kurtosis value is less than or equal to the preset value, kernel adaptive filtering is not performed, and the residual signal is directly used as the pure speech estimate.

9. A dot-matrix smart pen with AI recording function according to claim 8, characterized in that: The step of performing kernel adaptive filtering based on random Fourier features to output a clean speech estimate includes the following steps: The input vector is formed by taking the number of sampling points of the current filter order of the residual signal, and then mapping the input vector into a high-dimensional random feature vector through a pre-stored random Fourier feature mapping matrix. The kernel filter output is obtained by multiplying the high-dimensional random feature vector with the current kernel filter weight vector. The kernel filter output is then subtracted from the residual signal to obtain the clean speech estimate, and the kernel filter weight vector is updated.

10. A dot-matrix smart pen with AI recording function according to claim 9, characterized in that: The kernel width and step size factor of the kernel adaptive filter are adaptively adjusted according to dynamic parameters, including the following steps: The kernel width is determined based on the dynamic parameters, and the normalization step size factor is calculated based on the dynamic parameters. Based on the determined kernel width and step size factor, high-dimensional random feature vectors are generated and kernel filter weight vectors are updated.