Dual-branch rhythmogram fusion identification method and device, electronic equipment and storage medium

By using a dual-branch rhythm graph fusion identification method, the problems of large model parameters and low generalization ability in health monitoring are solved, realizing low-power real-time monitoring and reliable decision support, and improving the accuracy of health event identification and the immediacy of decision-making.

CN121564496BActive Publication Date: 2026-04-10QUANZHOU INST OF INFORMATION ENG
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-23
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies rely on subjective experience in health monitoring, which is difficult to scale. Deep learning solutions have large model parameters, cannot be embedded in low-power wearable devices, and have low generalization ability across devices and environments, failing to provide immediate and reliable decision-making basis.

Method used

A dual-branch rhythm graph fusion recognition method is adopted. Through preprocessing, overlapping segmentation, dual-branch complementary feature extraction and fusion, combined with a lightweight network architecture, real-time classification and decision support alerts for vibration and acoustic signals are achieved.

Benefits of technology

It improves the classification accuracy across datasets, transforms data into structured health event records in real time, generates reliable decision support alerts, and achieves low-power real-time monitoring, overcoming the problems of large size, poor generalization, and insufficient real-time performance of traditional models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564496B_ABST
    Figure CN121564496B_ABST
Patent Text Reader

Abstract

The application provides a dual-branch rhythmogram fusion recognition method and device, an electronic device and a storage medium. The method comprises: dividing a vibration sound signal to be recognized according to a preset time length and a preset overlap rate to obtain a plurality of vibration sound frames; converting each vibration sound frame into a rhythmogram; sequentially performing a multi-group convolution pooling operation with an increasing number of channels for each rhythmogram to obtain a first feature map corresponding to each vibration sound frame; sequentially performing a convolution operation and a convolution pooling operation with a decreasing and then increasing number of channels for each first feature map to obtain a second feature map corresponding to each vibration sound frame; fusing the first feature map and the second feature map corresponding to the same vibration sound frame to obtain a fusion feature corresponding to each vibration sound frame; outputting a classification result of each vibration sound frame based on the fusion feature corresponding to each vibration sound frame; converting the classification result into a health event record according to a preset rule, and generating a decision support reminder when a confidence threshold is met. The application has both accurate recognition and reliable decision-making capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of health monitoring, in particular to a dual-branch rhythmogram fusion recognition method and device, an electronic device and a storage medium. BACKGROUND

[0002] Vibroacoustic signals are one of the acoustic events reflecting the physiological state of the human body, which can be used for abnormal sound reminders in health monitoring scenarios. Traditional monitoring relies on the subjective experience of monitoring personnel, which has large subjective differences and poor consistency, and is difficult to scale. Although existing deep learning solutions can automatically identify, they generally have large model parameter quantities, cannot be embedded in low-power wearable devices, and have low cross-device and cross-environment generalization capabilities. Furthermore, even if the classification results are given, it is difficult to quickly determine a reliable decision basis according to the classification results. SUMMARY

[0003] The present application provides a dual-branch rhythmogram fusion recognition method and device, an electronic device and a storage medium, aiming to retain the advantages of lightweight dual-branch rhythmogram fusion recognition while converting classification results into health event records in real time, automatically generating decision support reminders after reliability detection, and solving the pain points of existing technologies that cannot provide reliable decision basis in time.

[0004] In a first aspect, an embodiment of the present application provides a dual-branch rhythmogram fusion recognition method, which includes: dividing a vibroacoustic signal to be recognized according to a preset time length and a preset overlap rate to obtain a plurality of vibroacoustic frames; converting each vibroacoustic frame into a rhythmogram to obtain a plurality of rhythmograms; sequentially performing multi-group convolution pooling operations with increasing channel numbers for each rhythmogram to obtain a first feature map corresponding to each vibroacoustic frame; sequentially performing convolution operations and convolution pooling operations with decreasing and then increasing channel numbers for each first feature map to obtain a second feature map corresponding to each vibroacoustic frame; fusing the first feature map and the second feature map corresponding to the same vibroacoustic frame to obtain a fusion feature corresponding to each vibroacoustic frame; outputting a classification result of each vibroacoustic frame based on the fusion feature corresponding to each vibroacoustic frame, wherein the classification result is used for health event reminders; converting the classification result into a health event record according to a preset rule; and generating a decision support reminder when the health event record meets a confidence threshold.

[0005] In a second aspect, the embodiments of the present application provide a dual-branch rhythmogram fusion recognition device, which comprises a segmentation module, a time-frequency conversion module, a first feature extraction module, a second feature extraction module, a fusion module, a classification output module, a classification conversion module, and a reminder generation module. The segmentation module is configured to segment a to-be-identified vibroacoustic signal according to a preset time length and a preset overlap rate to obtain a plurality of vibroacoustic frames. The time-frequency conversion module is configured to convert each vibroacoustic frame into a rhythmogram to obtain a plurality of rhythmograms. The first feature extraction module is configured to sequentially perform a multi-group convolution pooling operation with an increasing number of channels for each rhythmogram to obtain a first feature map corresponding to each vibroacoustic frame. The second feature extraction module is configured to sequentially perform a convolution operation and a convolution pooling operation with a decreasing and then increasing number of channels for each first feature map to obtain a second feature map corresponding to each vibroacoustic frame. The fusion module is configured to fuse the first feature map and the second feature map corresponding to the same vibroacoustic frame to obtain a fusion feature corresponding to each vibroacoustic frame. The classification output module is configured to output a classification result of each vibroacoustic frame based on the fusion feature corresponding to each vibroacoustic frame, and the classification result is used for a health event reminder. The classification conversion module is configured to convert the classification result into a health event record according to a preset rule. The reminder generation module is configured to generate a decision support reminder when the health event record meets a confidence threshold.

[0006] In a third aspect, the embodiments of the present application provide an electronic device, which comprises an electronic device main body and a master control device. The electronic device main body can be attached to or worn on the body surface of a human body to collect a to-be-identified vibroacoustic signal. The master control device is communicatively connected to the electronic device main body and comprises a memory and a processor. The memory is configured to store a computer program. The processor is configured to execute the computer program to implement the dual-branch rhythmogram fusion recognition method described above.

[0007] In a fourth aspect, the embodiments of the present application provide a computer readable storage medium for storing a computer program, which is executed to implement the dual-branch rhythmogram fusion recognition method described above.

[0008] The dual-branch rhythmogram fusion recognition method and device, the electronic device, and the storage medium described above can improve the cross-dataset classification accuracy to a daily monitoring threshold through preprocessing, overlapping segmentation, dual-branch complementary feature extraction and fusion, and overall cooperation of a lightweight network architecture, convert the classification result into a structured health event record in real time, automatically generate a decision support reminder after credibility detection, have both accurate recognition and reliable decision-making capabilities, and thus realize wearable rhythm monitoring in a low-power real-time operation mode, thereby overcoming the problems of a large volume, poor generalization, insufficient real-time performance, and inability to provide reliable decision-making basis of a traditional model. BRIEF DESCRIPTION OF DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from the drawings shown without creative labor.

[0010] Figure 1 The flow chart of the dual-branch rhythmogram fusion recognition method provided by the embodiments of the present application.

[0011] Figure 2 The flow chart of the step S101 sub-step provided by the embodiments of the present application.

[0012] Figure 3 The flow chart of the step S103 sub-step provided by the embodiments of the present application.

[0013] Figure 4 The flow chart of the step S104 sub-step provided by the embodiments of the present application.

[0014] Figure 5 The flow chart of the step S105 sub-step provided by the embodiments of the present application.

[0015] Figure 6 The structural block diagram of the dual-branch rhythmogram fusion recognition device provided by the embodiments of the present application.

[0016] Figure 7 The structural block diagram of the electronic device provided by the embodiments of the present application.

[0017] Figure 8 The schematic diagram of the electronic device provided by the embodiments of the present application.

[0018] Figure 9 The internal structure schematic diagram of the master control device applying the dual-branch rhythmogram fusion recognition method provided by the embodiments of the present application.

[0019] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the drawings. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the present application more clear, the following will further describe the present application with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0021] The terms "first", "second", "third", "fourth" and the like in the description and in the claims of the present application and above-mentioned drawings, if any, are used for distinguishing between similar objects, not necessarily described by their order or priority. It is to be understood that the data thus described can be interchanged, insofar as possible, without departing from the scope of the embodiments described. In other words, the described embodiments can be implemented according to an order other than the one illustrated or described herein. Furthermore, the terms "comprise" and "have" and any variations thereof, can also include other contents, for example, a process, a method, a system, a product or an apparatus comprising a series of steps or units does not necessarily have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or apparatuses.

[0022] It should be noted that the descriptions involving "first", "second" and the like in the present application are only for descriptive purposes, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" can explicitly or implicitly include one or more of the features. In addition, the technical solutions of various embodiments can be combined with each other, but must be based on the fact that a person skilled in the art can realize it. When the combination of technical solutions contradicts each other or cannot be realized, it should be considered that the combination of technical solutions does not exist and is not within the protection scope required by the present application.

[0023] Please refer to Figure 1 which is a flowchart of a dual-branch rhythmogram fusion recognition method provided by the embodiments of the present application. The present application provides a dual-branch rhythmogram fusion recognition method. Existing rhythm recognition usually uses a single-branch, large-parameter-depth model, which needs to rely on server-side operation, is difficult to run in real time on wearable electronic devices, and has limited cross-dataset generalization performance. The dual-branch rhythmogram fusion recognition method provided by the present application compresses the model volume while maintaining the cross-dataset accuracy of data features by overlapping segment segmentation, lightweight dual-feature fusion complementary features, and element-by-element fusion, realizing low-power, high-robust wearable rhythm real-time monitoring. The fusion recognition method comprises steps S101-S108.

[0024] In step S101, the vibration sound signal to be recognized is segmented according to a preset time length and a preset overlap rate to obtain a plurality of vibration sound frames.

[0025] In step S101, the vibration sound signal to be identified can be body surface microvibration, joint friction sound, respiratory sound, chewing sound or other original single-channel acoustic data collected by MEMS microphone, piezoelectric sheet, electronic stethoscope head, etc. The present application only takes heart microvibration as an example for illustration, and does not limit the signal source. Taking heart microvibration as an example, one complete activity period corresponds to one heart beat, mainly including the time duration of heart contraction and diastole, respectively called systole and diastole. Its typical vibration components usually include first heart sound (S1) and second heart sound (S2). The preset time length in the present application is generally 0.8-1.2s, so as to cover S1-S2 main components under adult 60-100bpm heart rate. The preset overlap rate can be set to 25-50%, so that the adjacent vibration sound frames obtained subsequently partially overlap in time domain, avoiding missing abnormal rhythm events at the boundary, and facilitating positioning the time of S1 and S2. Preferably, the preset time length is 0.8s, and the preset overlap rate is 50%.

[0026] Please refer to Figure 2 which is a flowchart of step S101 sub-step provided by the embodiment of the present application. The vibration sound signal to be identified is segmented according to the preset time length and the preset overlap rate, to obtain a plurality of vibration sound frames including steps S1011-S1012.

[0027] Step S1011, the vibration sound signal to be identified is obtained by sliding window interception to obtain a plurality of intercepted segments.

[0028] In step S1011, the window length of the sliding window is T, and the preset overlap rate is O, then the step length of the sliding window is Δ=T*(1-O), then the sliding window in the present application moves on the original vibration sound waveform with the window length T and the step length Δ. The sliding window moves once to intercept a continuous sampling point, to form a corresponding intercepted segment, to provide redundant and continuous time-frequency input for subsequent double feature fusion without increasing additional hardware cache, effectively improving the capture probability of occasional abnormal vibration. The abnormal vibration is used for health event reminder.

[0029] Step S1012, the intercepted segments with time length less than the preset time length are removed, and the remaining intercepted segments are retained to obtain a plurality of vibration sound frames.

[0030] In step S1012, each vibration sound frame contains at least one activity period. Since actual collection of the intercepted segment may produce tail short frame due to device start-stop or communication packet loss, the intercepted segment with window length less than T is filtered out by the preset time length, to ensure that each retained intercepted segment can cover at least one complete activity period, to lay a data foundation for double feature fusion to extract overall and detailed features with time sequence integrity. The activity period is a time segmentation reference of the vibration sound signal.

[0031] In the embodiment, before the vibration sound signal to be identified is divided according to the preset time length and the preset overlap rate, the method further includes the following steps: performing band-pass filtering on the vibration sound signal to be identified at a first preset frequency; resampling the vibration sound signal after the band-pass filtering at a second preset frequency; and performing amplitude normalization on the signal after the resampling to obtain the vibration sound signal to be identified. Specifically, the first preset frequency can be 25-400 Hz, which covers typical components of cardiac micro-vibration (such as S1 and S2) and other abnormal vibration energy concentration zones. The band-pass filtering can be Butterworth band-pass filtering, which can effectively suppress breathing, environmental low-frequency shaking and high-frequency electrical noise. The second preset frequency can be 2 kHz to retain sufficient spectral details. The amplitude normalization scales the signal to eliminate the gain difference of different MEMS microphones, piezoelectric sheets and other electronic device bodies used for collecting vibration sound signals, thereby ensuring the consistency of cross-device input and providing vibration sound signals to be identified with high signal-to-noise ratio and uniform dynamic range for double-feature fusion.

[0032] In step S102, each vibration sound frame is converted into a rhythm graph respectively to obtain a plurality of rhythm graphs.

[0033] In step S102, the rhythm graph can be a two-dimensional time-frequency matrix obtained by mapping each vibration sound frame to a mel scale through short-time Fourier transform (STFT). In this application, the size of the rhythm graph can be 256*256. The horizontal axis is the time axis, and the vertical axis is the 128-order mel scale frequency to cover the first preset frequency, which not only retains the rapid rising edge of typical mechanical vibration events (such as S1 and S2), but also compresses the image size, such as to 128*44 pixels. Then, the amplitude spectrum is logarithmically transformed and normalized to 0-255 gray levels to obtain a rhythm graph with a size of 256*256, which provides a representation basis for subsequent feature extraction that takes into account the temporal context and local spectral texture. In this application, S1 and S2 are exemplary mechanical vibration components.

[0034] In step S103, a plurality of convolution pooling operations with gradually increasing number of channels are sequentially performed on each rhythm graph to obtain a first feature map corresponding to each vibration sound frame.

[0035] In step S103, the rhythm graph is subjected to primary feature extraction through the plurality of convolution pooling operations with gradually increasing number of channels, so as to extract basic features such as edges and textures in the time-frequency domain through local receptive fields, significantly reduce the calculation complexity while retaining the main frequency characteristics, enable the complete transmission of intermediate features, and further extract global related features of the vibration sound signal to retain the time-frequency energy distribution pattern.

[0036] Please refer to Figure 3which is a flowchart of the step S103 sub-step provided by the embodiment of the present application. The multi-group convolutional pooling operation of gradually increasing the channel number is sequentially performed on each rhythm graph to obtain the first feature map corresponding to each vibration sound frame, including steps S1031-S1033.

[0037] In step S1031, the first group convolutional pooling operation with a channel number of 8 is performed on the rhythm graph to obtain a first sub-feature map.

[0038] In step S1031, the size of the rhythm graph is 1*256*256. A convolution kernel with a size of 3*3 can be used to expand the rhythm graph with a channel number of 8 from 1*256*256, while retaining the local texture on the mel frequency band; then, the maximum pooling with a size of 2*2 and a step of 2 is used to halve the height and width of the rhythm graph, while synchronously expanding the time-frequency receptive field with only a slight increase in parameters, effectively filtering out random noise and strengthening the typical mechanical vibration components (such as S1 and S2), thereby providing a stable bottom-level representation for subsequent hierarchical feature extraction.

[0039] In step S1032, the second group convolutional pooling operation with a channel number of 32 is performed on the first sub-feature map to obtain a second sub-feature map.

[0040] In step S1032, a convolution kernel with a size of 3*3 is still used, and the output channel is expanded from 8 to 32 to capture the cross-frequency harmonic coupling (such as the frequency multiplication component of abnormal vibration) on the first sub-feature map obtained in step S1031; then, the maximum pooling with a size of 2*2 and a step of 2 is still used to further compress the size of the first sub-feature map, and the time receptive field is also extended with only a slight increase in parameters, thereby refining the local time-frequency details and providing a compact and discriminative transition representation for subsequent global feature extraction.

[0041] In step S1033, the third group convolutional pooling operation with a channel number of 64 is performed on the second sub-feature map to obtain the first feature map.

[0042] In step S1033, a convolution kernel with a size of 3*3 is still used, and the channel number is doubled to 64 to fully fuse the cross-frequency association of S1 peak value and high-frequency abnormal vibration component on the second sub-feature map obtained in step S1032; then, the maximum pooling with a size of 2*2 and a step of 2 is still used to further compress the height and width of the second sub-feature map, so that the time receptive field is continuously expanded to cover a complete activity period, thereby converting the refined local time-frequency details into the first feature map representing the global time-frequency energy distribution, and providing a backbone representation with both semantics and compression ratio for subsequent fine-grained feature extraction.

[0043] In step S104, the convolution operation and the convolutional pooling operation with a gradually increasing and then decreasing channel number are sequentially performed on each first feature map to obtain the second feature map corresponding to each vibration sound frame.

[0044] In step S104, taking the first feature map obtained in step S103 as input, a deep and narrow channel feature refinement is performed by introducing a multi-scale feature fusion mechanism to perform fine-grained optimization on high-order features of the first feature map.

[0045] Please refer to Figure 4 , which is a flowchart of the sub-step S104 of step S104 provided by the embodiment of the present application. Convolution operation and channel number first decreasing and then increasing convolution pooling operation are sequentially performed on each first feature map to obtain the second feature map corresponding to each vibration and sound frame, including steps S1041-S1044.

[0046] In step S1041, the first convolution kernel is used to perform convolution and maximum pooling with a channel number of 32 on the first feature map to obtain a reduced dimension feature map.

[0047] In step S1041, the size of the first convolution kernel is consistent with the size of the convolution kernel used in the previous step S103, that is, 3x3. The first feature map with 64 channels is first compressed to 32 channels, and then the maximum pooling with a size of 2x2 and a step of 2 is used to reduce the height and width of the first feature map, which not only eliminates redundant high-frequency disturbances, but also retains abnormal low-frequency modulation information, laying a foundation for high signal-to-noise ratio for subsequent deep compression.

[0048] In step S1042, the first convolution kernel is used to perform convolution and maximum pooling with a channel number of 16 on the reduced dimension feature map to obtain a deep compressed feature map.

[0049] In step S1042, the first convolution kernel is used to reduce the channel number from 32 to 16, and at the same time, the maximum pooling with a size of 2x2 and a step of 2 is used to continue to reduce the height and width of the reduced dimension feature map, forcing the extraction of the most discriminative abnormal frequency, so that the difference between noise and normal spectrum is amplified to form a highly focused abnormal identifier.

[0050] In step S1043, the first convolution kernel is used to perform convolution and maximum pooling with a channel number of 8 on the deep compressed feature map to obtain a minimum channel feature map.

[0051] In step S1043, the minimum channel feature map represents the elimination of all redundancies and only retains the core frequency band strongly related to abnormal activity features, providing the most compact basis vector for subsequent channel recovery and reconstruction of high-frequency details.

[0052] In step S1044, the first convolution kernel is used to perform convolution and maximum pooling with a channel number of 16 on the minimum channel feature map to obtain a second feature map.

[0053] In step S1044, the channel is restored from 8 to 16 by using the first convolution kernel, and the channel expansion and the inverse compression of spatial invariance are realized by using the maximum pooling with the size of 2x2 and the step of 2, the compressed high-frequency components are reconstructed at a very low calculation amount, the sensitivity to short-time high-energy signals is enhanced, and the second feature map is output, which provides high-discrimination and low-redundancy abnormal description for the fusion decision.

[0054] In the embodiment, before the convolution and the maximum pooling of the first feature map by using the first convolution kernel with the channel number of 32, the method further includes the following steps: performing convolution processing on the first feature map by using a second convolution kernel to maintain the channel number of the first feature map and capture time-frequency correlation information in a preset time length range in the first feature map. Specifically, the size of the second convolution kernel is 5x5, and the second convolution kernel slides only on the time axis, and the channel number is maintained to be the channel number of the first feature map, that is, 64. The application can strengthen the sensitivity to short-time abnormal vibration and other correlation features without increasing the parameter burden by first establishing local time-frequency coupling and then performing dimension reduction in steps S1041-S1044, and can reserve more discriminative intermediate clues for the subsequent compression process. The abnormal frequency pattern and the abnormal identifier are vibration event features.

[0055] In step S105, the first feature map and the second feature map corresponding to the same vibration-sound frame are fused to obtain a fusion feature corresponding to each vibration-sound frame.

[0056] In step S105, the first feature map carries global time-frequency energy distribution, and the second feature map contains abnormal related high-frequency components, and the two are complementary in the channel dimension and the spatial dimension. The application superimposes the global time-frequency energy distribution and the abnormal related high-frequency components by element-by-element addition, so that each pixel retains the original energy information and obtains high-order discrimination enhancement, forming a unified representation with coarse-grained stability and fine-grained sensitivity.

[0057] Please refer to Figure 5 which is a flowchart of the sub-step of step S105 provided by the embodiment of the application. The first feature map and the second feature map corresponding to the same vibration-sound frame are fused to obtain a fusion feature corresponding to each vibration-sound frame, which includes steps S1051-S1053.

[0058] In step S1051, the first feature map is convolved and down-sampled to make the channel number of the first feature map consistent with that of the second feature map, and an aligned sub-map is obtained.

[0059] In step S1051, the first feature map with 64 channels is compressed to the same 16 channels as the second feature map by using a convolution kernel with a size of 1x1, the channel down-sampling is realized on the premise of maintaining the spatial resolution, the cross-channel information is reorganized, and the dimension mismatch in subsequent addition is avoided.

[0060] In step S1052, the aligned subgraph is maximum-pooled to make the spatial size of the aligned subgraph consistent with the second feature map, to obtain a spatially aligned feature map.

[0061] In step S1052, the aligned subgraph is maximum-pooled using a convolution kernel with a size of 2x2, to make the spatial size (i.e., height and width) of the aligned subgraph consistent with the size of the second feature map, so as to suppress potential noise peaks and ensure that the two features correspond to each other pixel by pixel during fusion.

[0062] In step S1053, the spatially aligned feature map and the second feature map are added element by element to obtain a fused feature.

[0063] In step S1053, the spatially aligned feature map and the second feature map are added element by element under the condition that they have the same tensor shape, so that the global time-frequency energy distribution and the abnormal-related high-frequency component are coupled point by point without additional parameters. After the element-by-element addition, the ReLU activation function is used for activation, so that the positions with high energy and high discrimination obtain stronger responses, and the low-value regions naturally decay, and finally the fused feature with stability and sensitivity is output.

[0064] In step S106, based on the fused feature corresponding to each vibrosonic frame, a classification result of each vibrosonic frame is output.

[0065] In step S106, the fused feature is input into a linear layer to output a classification result of each vibrosonic frame, such as normal, short-time abnormal vibration, diastolic abnormal vibration, and different categories. Specifically, the neurons are compressed to reduce redundancy, and then the Softmax is used to output the probability of each target category. In order to suppress the jitter between different vibrosonic frames, the probability average of the Softmax score of each vibrosonic frame is taken, and finally the confidence of the classification result of the whole vibrosonic signal is given. The confidence can be the score corresponding to different categories such as normal, short-time abnormal vibration, and diastolic abnormal vibration.

[0066] In step S107, the classification result is converted into a health event record according to a preset rule.

[0067] In step S107, the preset rule at least includes the following contents:

[0068] 1. Time window aggregation: the Softmax average score of a plurality of continuous vibrosonic frames in a corresponding sliding window is smoothed again to obtain a continuous confidence.

[0069] 2. Category threshold decision: if the continuous confidence of any category is greater than or equal to a first threshold value, and the number of continuous frames is greater than or equal to a preset frame number, the corresponding is marked as an abnormal event.

[0070] 3. Event encoding: write the results of the categorical thresholding decisions in a unified data structure, including event ID, device serial number, UTC timestamp, event category code, and continuous confidence;

[0071] 4. Confidence preliminary screening: discard the corresponding sliding window with continuous confidence less than the second threshold.

[0072] Generate a health event record according to the preset rules.

[0073] Step S108, when the health event record meets the confidence threshold, generate a decision support reminder.

[0074] In step S108, the decision support reminder can be received by the monitoring personnel to support the monitoring personnel to generate a corresponding decision according to the health event record. The present application can be divided into different levels of decision support reminders according to the difference between the continuous confidence of the health event record and the confidence threshold. For example, the decision support reminder can be a first-level reminder, a second-level reminder and a third-level reminder. When the decision support reminder is a first-level reminder, only the health event record is written back to the local log, and the monitoring personnel can check it when batch exporting, without real-time notification. When the decision support reminder is a second-level reminder, the health event record can be uploaded to a mobile applet, the applet writes in the user timeline in the background, and a silent message "rhythm change detected, please review later" is pushed to the monitoring personnel. When the decision support reminder is a third-level reminder, the mobile applet immediately forwards the health event record to the cloud health platform through an HTTPS encrypted channel, the platform completes fusion verification with the historical sequence to obtain a fusion confidence, so as to eliminate incidental noise and compensate for device drift, thereby outputting a more robust final confidence. The historical sequence is all health event records of the same user in the past time sorted by time, which is used as a personal baseline for horizontal comparison. If the fusion confidence is greater than the second threshold, the monitoring personnel is prompted to "review as soon as possible".

[0075] Please refer to Figure 6 which is a structural block diagram of a dual-branch rhythmogram fusion recognition device provided by an embodiment of the present application. The present application also provides a dual-branch rhythmogram fusion recognition device 10. The dual-branch rhythmogram fusion recognition device 10 includes a segmentation module 1, a time-frequency conversion module 2, a first feature extraction module 3, a second feature extraction module 4, a fusion module 5, a classification output module 6, a classification conversion module 7, and a reminder generation module 8.

[0076] The segmentation module 1 is used to segment the vibration sound signal to be recognized according to a preset time length and a preset overlap rate, to obtain a plurality of vibration sound frames.

[0077] The time-frequency conversion module 2 is used to convert each vibration sound frame into a rhythmogram, to obtain a plurality of rhythmograms.

[0078] The first feature extraction module 3 is configured to sequentially perform multi-group convolution pooling operations with increasing number of channels for each rhythmogram to obtain a first feature map corresponding to each vibrosonic frame.

[0079] The second feature extraction module 4 is configured to sequentially perform convolution operations and convolution pooling operations with decreasing and then increasing number of channels for each first feature map to obtain a second feature map corresponding to each vibrosonic frame.

[0080] The fusion module 5 is configured to fuse the first feature map and the second feature map corresponding to the same vibrosonic frame to obtain a fusion feature corresponding to each vibrosonic frame.

[0081] The classification output module 6 is configured to output a classification result of each vibrosonic frame based on the fusion feature corresponding to each vibrosonic frame.

[0082] The classification conversion module 7 is configured to convert the classification result into a health event record according to a preset rule.

[0083] The reminder generation module 8 is configured to generate a decision support reminder when the health event record meets a confidence threshold.

[0084] Referring to Figure 7 and Figure 8 , the present application further provides an electronic device 1000. The electronic device 1000 comprises an electronic device body 200 and a master control device 100. The electronic device body 200 is attached to or worn on the surface of a human body to collect vibrosonic signals to be identified. The master control device 100 is communicatively connected to the electronic device body 200 to complete rhythm identification and output a health reminder. In the present application, the electronic device 1000 can be a smart watch, a smart bracelet, a health tracking chest band, or a portable vibration recorder, etc. The electronic device body 200 is located on the back of the watchband, the inner side of the chest band, or the skin patch to realize daily collection.

[0085] Referring to Figure 9 , which is an internal structure diagram of the master control device for applying the double-branch rhythmogram fusion identification method provided by the embodiments of the present application.

[0086] As shown in Figure 9 , the master control device 100 comprises a memory 901 and a processor 902. The processor 902 is configured to run computer program instructions in the memory 901 to implement the double-branch rhythmogram fusion identification method.

[0087] The memory 901 includes at least one type of readable storage medium, including a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or a DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. The memory 901 can be an internal storage unit of the computer device in some embodiments, such as a hard disk of the computer device. The memory 901 can also be a storage device of an external computer device in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. configured in the computer device. Further, the memory 901 can include both an internal storage unit and an external storage device of the computer device. The memory 901 can be used to store application software installed in the computer device and various data, such as codes of the dual-branch rhythmogram fusion recognition method, and can also be used to temporarily store data that has been output or will be output.

[0088] Further, the host device 100 can also include a bus 903. The bus 903 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9 Only one thick line is used in the figure to represent the bus, but it does not mean that there is only one bus or only one type of bus.

[0089] Further, the host device 100 can also include a display component 904. The display component 904 can be an LED display, a liquid crystal display, a touch liquid crystal display, an organic light-emitting diode (OLED) touch, etc. The display component 904 can also be appropriately referred to as a display device or a display unit, and is used to display information processed in the host device 100 and to display a visualized user interface.

[0090] Further, the host device 100 can also include a communication component 905. The communication component 905 can optionally include a wired communication component and / or a wireless communication component (such as a Wi-Fi communication component, a Bluetooth communication component, etc.), and is usually used to establish a communication connection between the host device 100 and other computer devices.

[0091] Figure 9 Only the host device 100 with some components and implementing the dual-branch rhythmogram fusion recognition method is shown, and those skilled in the art can understand that,Figure 9 The illustrated structure does not constitute a limitation on the host device 100, and can include fewer or more components than shown, or combine certain components, or arrange the components differently.

[0092] In the above-described embodiments, the system, device and unit can be implemented wholly or partially by software, hardware or any combination thereof. When implemented by software, the system, device and unit can be implemented in the form of a computer program product wholly or partially.

[0093] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are wholly or partially generated. The computer device can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another via wire (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be stored by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, Solid State Disk (SSD)) and the like.

[0094] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.

[0095] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the above-described device embodiments are only schematic, for example, the division of the unit is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0096] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e., they may be located in one place, or they may be distributed on multiple network units. Part or all of the units may be selected according to actual needs to achieve the purpose of the embodiment.

[0097] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist independently, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0098] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only storage medium (ROM, Read-Only Memory), a random access storage medium (RAM, Random Access Memory), a magnetic disk or an optical disk, and various storage medium that can store program codes.

[0099] In the above embodiments, by pre-processing, overlapping segmentation, double-branch complementary feature extraction and fusion, and the overall cooperation of the lightweight network architecture, the cross-dataset classification accuracy is improved to the daily monitoring threshold, the classification result is immediately converted into a structured health event record, and after credibility detection, a decision support reminder is automatically generated, which has the dual ability of accurate identification and reliable decision, thereby realizing wearable rhythm monitoring in a low-power real-time running mode, overcoming the problems of large volume, poor generalization, insufficient real-time performance and inability to provide reliable decision basis of traditional models.

[0100] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application also intends to include these modifications and variations.

[0101] It should be understood that although the steps in the flowcharts of the drawings are shown in a sequential order following the arrows, the steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated otherwise, the execution of the steps is not strictly limited in order, and the steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order is not necessarily sequential, but can be round-robin or alternately executed with at least some of the other steps or sub-steps or stages of other steps.

[0102] The above merely shows the preferred embodiments of the present application, and of course cannot limit the scope of the right of the present application, so the equivalent changes made according to the claims of the present application still belong to the scope covered by the present application.

Claims

1. A dual-branch rhythmogram fusion recognition method, characterized in that, The dual-branch rhythmogram fusion recognition method comprises: segmenting the to-be-recognized vibroacoustic signal according to a preset time length and a preset overlap rate to obtain a plurality of vibroacoustic frames; converting each vibroacoustic frame into a rhythmogram respectively to obtain a plurality of rhythmograms; performing a plurality of convolution pooling operations with an increasing number of channels in each group on each rhythmogram in turn to obtain a first feature map corresponding to each vibroacoustic frame; performing convolution operations and convolution pooling operations with a decreasing number of channels in each group on each first feature map in turn to obtain a second feature map corresponding to each vibroacoustic frame; fusing the first feature map and the second feature map corresponding to the same vibroacoustic frame to obtain a fusion feature corresponding to each vibroacoustic frame; outputting a classification result of each vibroacoustic frame based on the fusion feature corresponding to each vibroacoustic frame; converting the classification result into a health event record according to a preset rule; generating a decision support reminder when the health event record meets a confidence threshold; wherein performing a plurality of convolution pooling operations with an increasing number of channels in each group on each rhythmogram in turn to obtain a first feature map corresponding to each vibroacoustic frame comprises: performing a first group of convolution pooling operations with a channel number of 8 on the rhythmogram to obtain a first sub-feature map; performing a second group of convolution pooling operations with a channel number of 32 on the first sub-feature map to obtain a second sub-feature map; performing a third group of convolution pooling operations with a channel number of 64 on the second sub-feature map to obtain the first feature map; wherein performing convolution operations and convolution pooling operations with a decreasing number of channels in each group on each first feature map in turn to obtain a second feature map corresponding to each vibroacoustic frame comprises: performing convolution and maximum pooling with a channel number of 32 on the first feature map using a first convolution kernel to obtain a dimension-reduced feature map; performing convolution and maximum pooling with a channel number of 16 on the dimension-reduced feature map using the first convolution kernel to obtain a deep compressed feature map; performing convolution and maximum pooling with a channel number of 8 on the deep compressed feature map using the first convolution kernel to obtain a minimum channel feature map; performing convolution and maximum pooling with a channel number of 16 on the minimum channel feature map using the first convolution kernel to obtain the second feature map; wherein fusing the first feature map and the second feature map corresponding to the same vibroacoustic frame to obtain a fusion feature corresponding to each vibroacoustic frame comprises: performing convolution down-sampling on the first feature map to make the channel numbers of the first feature map and the second feature map consistent to obtain an aligned sub-map; performing maximum pooling on the aligned sub-map to make the spatial dimensions of the aligned sub-map and the second feature map consistent to obtain a spatially aligned feature map; element-wise adding the spatially aligned feature map and the second feature map to obtain the fusion feature; wherein the plurality of vibroacoustic frames are obtained by sliding window truncation on the to-be-recognized vibroacoustic signal; and the preset rule comprises: re-smoothing the Softmax average scores of a plurality of continuous vibroacoustic frames in a corresponding sliding window to obtain continuous confidence; performing a category threshold decision; wherein if the continuous confidence of any category is greater than or equal to a first threshold and the number of continuous frames is greater than or equal to a preset frame number, the corresponding category is marked as an abnormal event. write the results of the category threshold decision in a uniform data structure, including event ID, device serial number, UTC timestamp, event category code, and continuous confidence; discard the corresponding sliding window whose continuous confidence is less than the second threshold.

2. The dual-branch rhythmogram fusion recognition method of claim 1, wherein, Before the vibration sound signal to be identified is segmented according to a preset time length and a preset overlap rate, the method further comprises: band-pass filtering the vibration sound signal to be identified at a first preset frequency; resampling the vibration sound signal after the band-pass filtering at a second preset frequency; amplitude normalization is performed on the resampled signal to obtain the vibration sound signal to be identified.

3. The dual-branch rhythmogram fusion identification method of claim 1, wherein, segmenting the vibration sound signal to be identified according to a preset time length and a preset overlap rate to obtain a plurality of vibration sound frames, comprising: sliding window interception is performed on the vibration sound signal to be identified to obtain a plurality of intercepted segments; removing the intercepted segments with a time length less than the preset time length, and retaining the remaining intercepted segments to obtain the plurality of vibration sound frames, wherein each vibration sound frame contains at least one active period.

4. The dual-branch rhythmogram fusion identification method of claim 1, wherein, Before the first feature map is convolved with a channel number of 32 using the first convolution kernel and maximum pooling, the method further comprises: convolving the first feature map using a second convolution kernel to maintain the channel number of the first feature map and capture the time-frequency correlation information in the first feature map within a preset time length range.

5. A dual-branch rhythmogram fusion recognition apparatus, characterized by, The dual-branch rhythmogram fusion recognition device comprises: a segmentation module configured to segment the vibration sound signal to be identified according to a preset time length and a preset overlap rate to obtain a plurality of vibration sound frames; a time-frequency conversion module configured to convert each vibration sound frame into a rhythmogram to obtain a plurality of rhythmograms; a first feature extraction module configured to sequentially perform a plurality of convolution and pooling operations with an increasing number of channels for each rhythmogram to obtain a first feature map corresponding to each vibration sound frame; a second feature extraction module configured to sequentially perform convolution operations and convolution and pooling operations with a decreasing and then increasing number of channels for each first feature map to obtain a second feature map corresponding to each vibration sound frame; a fusion module configured to fuse the first feature map and the second feature map corresponding to the same vibration sound frame to obtain a fusion feature corresponding to each vibration sound frame; a classification output module configured to output a classification result of each vibration sound frame based on the fusion feature corresponding to each vibration sound frame, wherein the classification result is used for health event reminding; a classification conversion module configured to convert the classification result into a health event record according to a preset rule; a reminding generation module configured to generate a decision support reminder when the health event record meets a confidence threshold. The method further comprises: performing a first group of convolution and pooling operations with a channel number of 8 on the rhythmogram to obtain a first sub-feature map; performing a second group of convolution and pooling operations with a channel number of 32 on the first sub-feature map to obtain a second sub-feature map; performing a third group of convolution and pooling operations with a channel number of 64 on the second sub-feature map to obtain the first feature map; The method further comprises: performing convolution operations and convolution and pooling operations with a decreasing and then increasing number of channels for each first feature map to obtain a second feature map corresponding to each vibration sound frame, comprising: performing a first group of convolution and pooling operations with a channel number of 8 on the rhythmogram to obtain a first sub-feature map; performing a second group of convolution and pooling operations with a channel number of 32 on the first sub-feature map to obtain a second sub-feature map; performing a third group of convolution and pooling operations with a channel number of 64 on the second sub-feature map to obtain the first feature map; performing convolution operations and convolution and pooling operations with a decreasing and then increasing number of channels for each first feature map to obtain a second feature map corresponding to each vibration sound frame, comprising: performing a first group of convolution and pooling operations with a channel number of 8 on the rhythmogram to obtain a first sub-feature map; performing a second group of convolution and pooling operations with a channel number of 32 on the first sub-feature map to obtain a second sub-feature map; performing a third group of convolution and pooling operations with a channel number of 64 on the second sub-feature map to obtain the first feature map; performing channel number 32 convolution and maximum pooling on the first feature map by using the first convolution kernel to obtain a reduced dimension feature map; performing channel number 16 convolution and maximum pooling on the reduced dimension feature map by using the first convolution kernel to obtain a deep compressed feature map; performing channel number 8 convolution and maximum pooling on the deep compressed feature map by using the first convolution kernel to obtain a minimum channel feature map; performing channel number 16 convolution and maximum pooling on the minimum channel feature map by using the first convolution kernel to obtain the second feature map; wherein the first feature map and the second feature map corresponding to the same vibroacoustic frame are fused to obtain a fusion feature corresponding to each vibroacoustic frame, comprising: performing convolution downsampling on the first feature map to make the channel numbers of the first feature map and the second feature map consistent to obtain an aligned subgraph; performing maximum pooling on the aligned subgraph to make the spatial dimensions of the aligned subgraph and the second feature map consistent to obtain a spatially aligned feature map; element-wise adding the spatially aligned feature map and the second feature map to obtain the fusion feature; wherein the plurality of vibroacoustic frames are obtained by sliding window cutting on the to-be-identified vibroacoustic signal; the preset rule comprises: performing re-smoothing on the Softmax average score of the continuous plurality of vibroacoustic frames in the corresponding sliding window to obtain a continuous confidence; performing a class threshold decision; wherein if the continuous confidence of any class is greater than or equal to a first threshold value, and the number of continuous frames is greater than or equal to a preset frame number, then the corresponding is marked as an abnormal event; writing each result of the class threshold decision in a unified data structure, including event ID, device serial number, UTC timestamp, event class code and continuous confidence; discarding the corresponding sliding window whose continuous confidence is less than a second threshold value.

6. An electronic device, comprising: The electronic device comprises: an electronic device body, which can be attached or worn on the surface of the human body, and is used to collect a to-be-identified vibroacoustic signal; and a master control device, which is in communication connection with the electronic device body, and comprises: a memory, which is used to store a computer program; and a processor, which is used to execute the computer program to realize the dual-branch rhythmogram fusion recognition method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer readable storage medium is used to store a computer program, and the computer program is executed to realize the dual-branch rhythmogram fusion recognition method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Mel spectrum and lightweight neural network-based debris flow infrasound identification method

    CN120823850A

  • Optical fiber pickup mode recognition method and device, electronic equipment and storage medium

    CN121278482A