Multimodal Data Perception and Processing Method and System for Audio-Intelligent Film
By constructing a model state transition diagram and spatial position matrix and dynamically adjusting the processing strategy, the problems of abnormal identification of multimodal data processing and accurate identification of interference areas in the sound and intelligence membrane structure are solved, efficient data processing and resource utilization are achieved, and the robustness and stability of the system are improved.
Patent Information
- Application Number
- CN202510764828.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing multimodal data processing methods are difficult to maintain effective performance in the dynamic and local interference-frequency environment in the sound and intelligence membrane structure, resulting in abnormal data misjudgment and processing of flow chain errors, and cannot meet the dual requirements of accuracy and stability.
Construct a model state transition diagram and spatial position matrix, and calibrate the cache nodes and interference areas of the abnormal data, combine the modal compression strategy and the delay decoding mechanism to dynamically adjust the processing strategy to achieve accurate identification and processing of abnormal modal data.
It improves the system's response speed and processing flexibility to abnormal mode data, reduces the burden of data transmission and calculation, ensures the continuity and stability of the main processing flow, and improves resource utilization efficiency and robustness.
Smart Images

Figure CN120281754B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a multimodal data perception and processing method and system for a sound-intelligent membrane. Background Art
[0002] In the process of multimodal data processing, integrating heterogeneous modal data such as images, audio, temperature, humidity, and vibration to form a collaborative perception and dynamic response mechanism is a key issue in current intelligent system design. In recent years, with the development of acoustic control technology and soft structural materials, air membrane structures (i.e., sound-intelligent membranes) that integrate acoustic functions with intelligent perception have gradually become important components in new scalable spaces. They offer advantages such as light weight, flexible deployment, and high functional integration.
[0003] In this context, embedding multimodal data processing strategies into sound-intelligence membranes to achieve adaptive recognition and response to environmental conditions, human activity, and noise sources has become a research hotspot. However, because sound-intelligence membrane structures are often located in semi-enclosed or highly reverberant environments, perception signals are often accompanied by non-stationary interference, spatial interference overlap, and modal information drift. This makes it difficult for existing perception processing methods to meet the dual requirements of accuracy and stability.
[0004] Existing multimodal perception technologies often employ parallel modeling, attention fusion, or data-level alignment to improve the compatibility and information integrity of data from different modalities. However, these methods often rely on stable signal input conditions and struggle to maintain effective performance in dynamic environments with frequent local interference. For example, in a sound-intelligence membrane structure, because audio and image modalities are limited by physical propagation paths and structural occlusions, abnormal data in local areas can easily lead to misjudgment, which in turn triggers chain errors in the processing flow. Summary of the Invention
[0005] In view of the problems existing in the above background technology, the present invention is proposed.
[0006] Therefore, the problem to be solved by the present invention is how to construct a multimodal perception framework that can transfer states over time, support abnormal node marking, and have location association and compression selection processing capabilities under the premise of ensuring information integrity.
[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0008] In the first aspect, the present invention provides a multimodal data perception and processing method for a sound-intelligence membrane, which includes constructing a modal state transition diagram based on the collected multimodal data, and calibrating the cache nodes of abnormal data in the state transition path; constructing a spatial position matrix based on the spatial position relationship of the sensor array in the sound-intelligence membrane structure, and determining the interference areas in the audio and image modalities and marking the interference modal labels through the partition numbers and strong interference identification thresholds in the spatial position matrix; when the modal type of a partition in the interference modal label table is marked as an interference state, or the corresponding node in the modal state transition diagram has been marked as an abnormal cache node, the modal data involved is switched to a compressed processing state according to a preset modal compression strategy rule; combined with the compression state mark and the spatial position matrix index of the corresponding position, the modal data in the cache node is selectively decoded or discarded according to the set delay, and backfilled to the main processing flow according to the state transition path.
[0009] As a preferred solution of the multimodal data perception and processing method for the sound intelligent membrane described in the present invention, wherein: the calibration of the cache node includes: extracting the main feature parameter sequence representing the state evolution from the multimodal; constructing the state encoding vector with the main feature combination of the multimodal at the current moment , and construct a state coding sequence; convert the state coding sequence into a node state sequence, construct directed edges in the time direction, and form a modal state transition graph, where each node represents a state coding vector and each edge represents a state transition under a timestamp; perform abnormal state transition detection in the modal state transition graph path and calibrate abnormal cache nodes.
[0010] As a preferred solution of the multimodal data perception and processing method for the sound intelligence membrane of the present invention, the formula for abnormal state transition detection is as follows:
[0011] ,
[0012] in, is the L1 norm, and represents the modal state vector at two consecutive time points, Indicates time arrive The change amplitude of the state encoding vector between , then the modal data corresponding to the current time point is cached as an abnormal candidate data block.
[0013] As a preferred solution of the multimodal data perception and processing method for the sound-intelligence membrane described in the present invention, the construction of the spatial position matrix includes: based on the actual layout coordinate position of each modal sensor in the sound-intelligence membrane structure, extracting the geometric coordinate information of all sensors in the two-dimensional membrane surface, and constructing an initial coordinate set , where the initial coordinate set Each element of consists of a sensor number and a position pair in the membrane surface coordinate system (X, Y); the initial coordinate set Map rows and columns in numerical order, divide the two-dimensional membrane surface into several spatial partitions, and based on the initial coordinate set in each partition Element aggregation generates a spatial position matrix , where the spatial position matrix Each element in contains the partition number and the number of modal sensors distributed in the corresponding block.
[0014] As a preferred embodiment of the multimodal data perception processing method for the sound intelligent membrane of the present invention, wherein: the interference area in the audio and image modes is determined by: A set of all partitions containing multi-source modal sensors is extracted from the dataset, and a multimodal interference candidate partition set is constructed by combining the node state indicators in the modal state transition diagram. The interference intensity index is determined for each partition in the multimodal interference candidate partition set, and the partition whose interference intensity index exceeds the interference determination threshold is marked as an interference area by setting the interference determination threshold.
[0015] As a preferred embodiment of the multimodal data perception and processing method for the sound intelligent membrane described in the present invention, the process of determining the interference intensity index includes: performing time domain feature analysis on the audio modal data segments within each spatial partition to extract the audio waveform disturbance level parameters; performing inter-frame pixel fluctuation density analysis on the image modal data of each spatial partition to form image level parameters; and combining the audio waveform disturbance level parameters and the image level parameters to form the interference intensity index.
[0016] As a preferred embodiment of the multimodal data perception and processing method for the sound intelligence membrane described in the present invention, the switching of the modal data involved to the compression processing state includes: screening out a set of cache node numbers that have been marked as abnormal states according to the partition number and modal type in each record in the interference modal label table, combined with the state encoding vector of the corresponding node in the modal state transition diagram; establishing a compression state switching mapping table based on the cache node number set; fusing the compression state switching mapping table with the modal state transition diagram, positioning and updating according to the node number, marking the compression state flag in the modal state transition diagram, and appending the corresponding compression strategy number to form an updated modal state transition diagram; reading each abnormal node number that has been set to the compression state in turn, comparing it with the modal data stored in the cache node, selecting the corresponding compression structure according to the modal type, compressing it, and synchronously updating the cache node.
[0017] As a preferred solution of the multimodal data perception and processing method for the sound intelligent membrane described in the present invention, wherein: the selective decoding or discarding of the modal data in the cache node according to the set delay includes: retrieving the modal type set currently in the compressed state in the compression state switching table, and extracting the matching maximum tolerable delay parameter according to the corresponding compression strategy ID, and constructing a compression mode-delay threshold mapping table; according to the spatial position matrix, combined with the cache node number set, each cache node is numbered, and the coordinate index in the spatial topology is located by the matrix index to obtain the modal data set of the corresponding cache area; for each data record in the modal data set , extract the timestamp and compare it with the current system time. If the timestamp is less than the maximum delay threshold corresponding to the current system time compression mode-delay threshold mapping table, the corresponding data will be marked as expired and added to the pending cache list; for each cache record in the pending cache list, according to the update mode state transition diagram, the path priority score of the corresponding node is extracted and compared with the median of the path priority list: when the priority score is less than the median of the path priority list, it is considered that the importance of the path is not enough to support the delayed decoding cost, and the processing flag is set to be directly discarded; otherwise, it is set to be delayed and retained, and enters the decoding queue.
[0018] As a preferred solution of the multimodal data perception and processing method for the sound intelligence membrane described in the present invention, the backfilling to the main processing flow includes: for the modal data entries that have been set to delayed retention, calling the decoder corresponding to the compression strategy ID to decode the compressed format back to the original modal structure format; for each decoded modal data, further reading the original path mark saved in the modal state transition diagram as the backfill target path number; the modal data that has been decoded and marked with the path number is sent to the main processing flow inlet channel in the order of the path number, and the delayed cache data source tag is added at the same time.
[0019] In a second aspect, the present invention provides a multimodal data perception and processing system for a sound intelligence membrane, comprising:
[0020] The modal diagram construction module constructs a modal state transition diagram based on the collected multimodal data and calibrates the cache nodes of abnormal data in the state transition path; the spatial encoding module constructs a spatial position matrix based on the spatial position relationship of the sensor array in the sound intelligence membrane structure, and determines the interference area in the audio and image modalities and marks the interference modal label through the partition number and strong interference identification threshold in the spatial position matrix; the modal compression module, when the modal type of a partition in the interference modal label table is marked as an interference state, or the corresponding node in the modal state transition diagram has been marked as an abnormal cache node, the modal data involved is switched to the compression processing state according to the preset modal compression strategy rules; the delay decoding module combines the compression state mark and the spatial position matrix index of the corresponding position to selectively decode or discard the modal data in the cache node according to the set delay, and backfill it to the main processing flow according to the state transition path.
[0021] The beneficial effects of the present invention are as follows: by constructing a modal state transition diagram and a spatial position matrix, the present invention can accurately identify abnormal states and interference areas, and dynamically adjust the processing strategy, thereby effectively improving the system's response speed and processing flexibility to abnormal modal data. Especially in the case of frequent interference data or limited bandwidth resources, the present invention greatly reduces the data transmission and computing burden through modal compression and delayed decoding mechanisms, ensuring the continuity and stability of the main processing flow. At the same time, the method has strong adaptive capabilities, and can dynamically adjust cache management and path priority according to real-time state changes, effectively preventing data congestion and delay diffusion. In summary, the present invention significantly improves the resource utilization efficiency and robustness of the multimodal processing system while ensuring data perception accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 The flowchart of the multimodal data perception and processing method for the sound intelligent membrane.
[0024] Figure 2 This is a structural diagram of the multimodal data perception and processing system for the sound intelligence membrane. DETAILED DESCRIPTION
[0025] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0026] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0027] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0028] As mentioned in the above background technology, embedding multimodal data processing strategies into sound-intelligence membranes to achieve adaptive recognition and response to environmental conditions, human activities and noise sources has become a research hotspot. However, since sound-intelligence membrane structures are mostly in semi-enclosed or high-reverberation environments, perception signals are often accompanied by non-stationary interference, spatial interference overlap, and modal information drift, making it difficult for existing perception processing methods to meet the dual requirements of accuracy and stability. In existing multimodal perception technologies, parallel modeling, attention mechanism fusion or data-level alignment are often used to improve the adaptability and information integrity between different modal data. However, such methods often rely on stable signal input conditions and are difficult to maintain effective performance in dynamic environments with frequent local interference. For example, in sound-intelligence membrane structures, since audio and image modalities are limited by physical propagation paths and structural occlusions, abnormal data in local areas are very likely to be misjudged, which in turn causes chain errors in the processing flow.
[0029] Figure 1 FIG is a flow chart of a multimodal data perception and processing method for a sound intelligent membrane according to an embodiment of the present invention. Figure 1 As shown, the multimodal data perception and processing method for the sound intelligence membrane includes:
[0030] S1: Construct a modal state transition diagram based on the collected multimodal data and calibrate the cache nodes of abnormal data in the state transition path.
[0031] S1.1: Extract the main feature parameter sequence representing the state evolution from the three modalities of audio, image, and vibration; construct the state encoding vector based on the main feature combination of the three modalities at the current moment , where each component is encoded into a finite set of states (e.g., “low, medium, high” correspond to 0 / 1 / 2) through a hierarchical mapping method (e.g., threshold, interval), so that the state encoding vector can be used to construct a state encoding sequence.
[0032] S1.2: Encode the state sequence It is converted into a node state sequence, and directed edges are constructed in the time direction to form a modal state transition graph, where each node represents a state encoding vector and each edge represents a state transition under a timestamp.
[0033] Among them, only the occurrence frequency of each transition path is retained in the modal state transition diagram As a basic attribute, a path priority list is constructed based on the path frequency. The higher the frequency, the higher the priority. This priority will be used as the basis for the path selection strategy in the interference status backfill.
[0034] S1.3: In the modal state transition diagram path, use the following abnormal state transition detection formula to calibrate the abnormal cache node:
[0035] ,
[0036] in, is the L1 norm, and represents the modal state vector at two consecutive time points, Indicates time arrive The change amplitude of the state encoding vector between (The threshold is determined by the historical average transition level of the training set), then the modal data corresponding to the current time point is cached as an abnormal candidate data block, and the node number and the mode of occurrence are recorded.
[0037] S2: Based on the spatial position relationship of the sensor array within the sound-intelligence membrane structure, a spatial position matrix is constructed. The interference areas in the audio and image modalities are determined and the interference modal labels are marked using the partition numbers and strong interference identification thresholds in the spatial position matrix.
[0038] S2.1: Construct a spatial position matrix.
[0039] Based on the actual layout coordinates of each modal sensor in the sound-intelligence membrane structure, the geometric coordinate information of all sensors in the two-dimensional membrane surface is extracted to construct the initial coordinate set. ,in Each element of consists of a sensor number and a position pair in the membrane surface coordinate system (X, Y).
[0040] The initial coordinate set By mapping rows and columns in numerical order, the two-dimensional membrane surface is divided into several spatial partitions using spatial clustering and region encoding algorithms. Several methods can be used for spatial region partitioning. The Voronoi diagram method is suitable for adaptive partition boundary partitioning in non-uniform distributions, ensuring that each partition surrounds the nearest neighboring sensor node as much as possible. DBSCAN, on the other hand, effectively generates distribution regions in scenarios with uneven density, facilitating the dynamic identification of abnormal cluster distribution areas. There is no single method to choose from; it can be flexibly selected based on the actual deployment situation.
[0041] Based on the initial coordinate set in each partition Element aggregation generates a spatial position matrix , where the spatial position matrix Each element in contains the partition number and the number of modal sensors distributed in the corresponding block.
[0042] Furthermore, if the physical location of abnormal nodes identified in the state transition diagram cannot be clearly determined, interference identification based on the spatial clustering effect cannot be achieved. To this end, the present invention introduces a joint spatial partition-modal state encoding field in each abnormal node data structure. This encoding field consists of three subfields: the spatial block number, the sensor modality type, and the corresponding state code.
[0043] S2.2: Determine the interfering regions in the audio and image modalities and label the interfering modalities.
[0044] After generating a spatial location matrix containing the distribution of modal types, the system begins screening candidate interference regions. The focus here is on identifying "multimodal coupling interference," where abnormal nodes exist in multiple modes in the same area.
[0045] In the operation, first in the spatial position matrix A set of all partitions containing multi-source modal sensors is extracted from the dataset, and a set of multi-modal interference candidate partitions is constructed by combining the node state indicators in the modal state transition graph, where each partition is a clustered area that contains abnormal state nodes at the same time in spatial position.
[0046] An interference strength index is determined for each partition in the multimodal interference candidate partition set, wherein the interference strength index is formed according to the joint disturbance of the waveform instability index in the audio modality and the pixel area noise dispersion in the image modality, and the partition whose interference strength index exceeds the medium interference area is marked as an interference area.
[0047] The mode types in all interference areas are marked in the interference mode label table, where each record in the interference mode label table contains the interference space partition number, mode type and interference intensity index, and the label table is written back to the node attributes of the original modal state transition diagram.
[0048] Preferably, the generation process of the interference intensity index includes:
[0049] The time domain feature analysis is performed on the audio modal data segment in each spatial partition to extract the audio waveform disturbance level parameters.
[0050] Specifically, the audio signal is first divided into multiple sub-segments according to the preset time window, and the stability determination process is performed for each sub-segment: if the amplitude of the continuous fluctuation of signal energy in any sub-segment exceeds twice the average of the entire time window, and the abnormal fluctuation lasts for more than three sub-segments, then the audio segment is marked as a strong disturbance segment; if there is only one abnormal fluctuation, and it lasts for less than two sub-segments, it is marked as a light disturbance segment; the rest are marked as medium disturbance segments. After counting the proportion of each type of disturbance segment, the ratio of strong disturbance segments is determined based on whether it exceeds U% of the total number of segments, where the U value can be used to determine the optimal threshold through the ROC curve to distinguish between interference and non-interference areas. Further classification is performed: if it exceeds, the audio disturbance level of the partition is determined as an audio high disturbance area; if the strong disturbance segment and the medium disturbance segment exceed 2U% in total, it is marked as an audio medium disturbance area; otherwise, it is an audio low disturbance area.
[0051] The inter-frame pixel fluctuation density analysis is performed on the image modal data of each spatial partition to form image grade parameters.
[0052] First, region matching and difference calculations are performed on consecutive image frames to extract the total area ratio of pixel grayscale variation in the target region between any three consecutive frames. If this ratio exceeds 20% of the entire image area three times in a row (this data is a preset value and should be verified in practice), it is considered a high-fluctuation segment. If it occurs discontinuously but more than once, it is considered a medium-fluctuation segment; the rest are low-fluctuation segments. Based on this, the distribution density and spatial concentration of each type of fluctuation segment are calculated. If a high-fluctuation segment appears in the center of the image region for more than five consecutive frames and covers more than h% of the central area (this data should be verified in practice), the image region is marked as a strongly disturbed area. If the high-fluctuation segment is mainly concentrated at the edge or only briefly jumps (for example, within three frames), it is marked as a slightly disturbed area. All other cases are classified as moderately disturbed areas.
[0053] After obtaining the audio and image disturbance levels, they are combined to generate a disturbance intensity index label for the current spatial partition according to the following rules: If the audio disturbance level is high and the image disturbance level is strong or medium, the current partition is labeled as a high-disturbance area. If the audio disturbance level is medium and the image disturbance level is strong or medium, the current partition is labeled as a medium-disturbance area. If the audio disturbance level is low but the image disturbance level is strong, the current partition is still labeled as a medium-disturbance area. Only when both levels are low is the current partition labeled as a non-disturbance area. Furthermore, if the image disturbance level is high but its fluctuations are confined to the edge and do not affect the recognition quality of the central part of the image, the original high-disturbance label is downgraded to medium-disturbance.
[0054] It should be noted that vibration modal characteristics are mainly used for state sequence modeling and path branch identification, and do not participate in the determination of disturbance intensity indicators.
[0055] Given the high overlap in input data between the modal state transition diagram construction process and the spatial matrix generation process, to improve processing efficiency, the operations in S1 and S2 can be executed simultaneously through a thread-level parallel mechanism. However, since both involve reading and writing node states and structure identifiers, to avoid data contention and state conflicts, this paper introduces a lightweight distributed lock mechanism for synchronization. The specific mechanism is as follows:
[0056] A state lock tag is attached to each node state record structure, and the S1 and S2 sub-threads apply for access rights before read and write operations; state write operations adopt a write-first strategy, and read operations adopt a lock snapshot mechanism to avoid inconsistent propagation; a timestamp-based distributed conflict mediation strategy is adopted to ensure that concurrent tasks reach consensus at both the spatial topology and state flow transition levels.
[0057] S3: When the modal type of a partition in the interference modal label table is marked as an interference state, or the corresponding node in the modal state transition diagram has been marked as an abnormal cache node, the modal data involved will be switched to the compression processing state according to the preset modal compression strategy rules.
[0058] When it is detected that any record in the interference modal label table indicates that a block has interference anomalies such as audio, image or vibration, or a node in the modal state transition diagram has shown a persistent abnormal cache state, it indicates that the modal data can no longer express information accurately according to the conventional path or has no redundant interference filtering capabilities. On this basis, the present invention designs a dual trigger mechanism based on "interference trigger + cache anomaly" to ensure a highly timely response to data anomalies in a multi-source modal system. Unlike traditional systems that often use modal anomalies as a post-processing stage, the present invention moves the compression processing forward to the data stream distribution center stage, effectively avoiding modal pollution and cache resource waste caused by abnormal propagation.
[0059] To achieve the above switching process, first, based on the partition number and modal type in each record in the interference modal label table, combined with the state encoding vector of the corresponding node in the modal state transition diagram, all cache node number sets that have been calibrated as abnormal states are screened out. This operation is essentially a cross-fusion of the data calibration results from two sources. On the one hand, it comes from the historical state evolution chain in the modal state transition diagram, and on the other hand, it comes from the interference label table jointly determined by the spatial position matrix and the modal recognition model in the previous stage. The combination of the two can form a set of highly reliable compressed candidate nodes. During the operation, a fast screening path is established through the hash mapping function, so that the compression recognition time of nodes in large-scale modal graphs does not exceed linear complexity, thereby enhancing processing efficiency.
[0060] Based on the cache node number set, a compression state switching mapping table is established. Each mapping record contains the abnormal node number, modality type, spatial partition number, and the proposed compression strategy number. It is worth emphasizing that this mapping table is not a static structure but is generated in real time based on dynamic variables such as the current operating load, adaptive network bandwidth, and cache availability to ensure appropriate matching of compression strategies. For example, under high concurrency interference conditions, the compression strategy number will favor a high compression ratio and high robustness structure. Under normal conditions, a lightweight strategy will be prioritized to maintain information fidelity.
[0061] The compression strategy number is selected based on preset modal compression strategy rules, matching the corresponding compression method to the modal type (audio, image, vibration). For audio modalities, priority is given to methods such as low-frequency information discarding, wavelet packet feature coding, and speech main contour extraction. Their core purpose is to reduce background noise interference and preserve the main signal form. For image modalities, compression methods such as regional pixel aggregation, edge sparse reconstruction, or low-resolution version replacement are used to reduce visual contamination caused by pixel interference. For vibration modalities, strategies such as main frequency band focusing, redundant period removal, and Fourier peak clipping are used to compress data from the spectral dimension. All strategies are equipped with a compression ratio selection, allowing the compression level to be automatically selected based on the current network load. For example, the product of the predicted transmission delay function and the cache read frequency is used as a dynamic reference for the compression factor.
[0062] Subsequently, the compression state switching mapping table is merged with the modal state transition diagram, and the location update is performed according to the node number. The compression state flag is marked in the modal state transition diagram, and the corresponding compression strategy number is attached, so that the abnormal nodes in the modal state transition diagram have a compression attribute field, forming an updated modal state transition diagram, which is convenient for subsequent cache scheduling and data reading modules to distinguish and process. The technical key to this operation is to maintain the topological integrity of the original state transition diagram while introducing a compression flag, so that the subsequent scheduler can clearly distinguish between regular nodes and compressed nodes when reading the graph. Unlike traditional systems that use additional compression channels for bypass processing, the present invention realizes unified management of data paths through embedded attribute updates in the graph structure, avoiding additional index overhead and concurrency consistency issues.
[0063] Furthermore, based on the records in the compression state switching mapping table, the number of each abnormal node that has been set to the compressed state is read in sequence. The modal data stored in the abnormal cache node is compared with the modal data. Based on the modal type, the corresponding compression structure (e.g., feature envelope for audio modal compression, regional pixel encoding for image modal compression, and main frequency parameter set for vibration modal compression) is selected for compression. The cache node is then updated synchronously, annotated with the compression structure type and compression version number.
[0064] When the compression structure is replaced, the compression status update is pushed to the modal scheduler. Based on this push information, the scheduler replans the node scheduling path and places the compression nodes in a medium-to-low priority level during the subsequent data reading and cache scheduling phases. This avoids frequent calls to compression nodes when resources are limited, thereby reducing the impact of redundant data on system performance.
[0065] In subsequent operations, if it is detected that the interference intensity is lower than the switchback threshold, the system automatically calls the preset compression state recovery mechanism, switches the node state back to the regular data processing path, and restores the original cache structure and scheduling logic.
[0066] S4: Combined with the compression state mark and the spatial position matrix index of the corresponding position, the modal data in the cache node is selectively decoded or discarded according to the set delay, and backfilled to the main processing flow according to the state transfer path.
[0067] After executing step S3 to complete the main processing flow and cache separation strategy, in order to achieve effective management and recovery of cached data, it is necessary to evaluate the delay tolerance and make decisions on whether to keep or remove the modal data in the buffer state based on the compression state mark and spatial position matrix index.
[0068] First, the set of modal types currently in the compressed state is retrieved from the compression state switching table, and the maximum tolerable delay parameter that matches each modal entry in the set is extracted based on the compression strategy ID corresponding to each modal entry in the set, thereby constructing a compression mode-delay threshold mapping table.
[0069] Next, based on the spatial position matrix constructed in step S2 and in combination with the cache node number set, for each cache node number, its coordinate index in the spatial topology is located through the matrix index, and then the modal data set of its corresponding cache area is obtained.
[0070] For each data record in the modal data set, extract its timestamp and compare it with the current system time. If the timestamp is less than the current system time minus the maximum delay threshold corresponding to the compression modality-delay threshold mapping table, indicating that the modal data has exceeded the maximum delay threshold tolerated by the compression strategy, the data is marked as "expired" and added to the pending cache list.
[0071] This marking process ensures that subsequent retention and removal operations are performed only on modal data that has actually timed out, avoiding unnecessary processing overhead on valid data.
[0072] Then, for each cache entry in the pending cache list, the path priority score of the corresponding node is extracted based on the updated modal state transition diagram constructed in step S3 and compared with the median of the path priority list. If the priority score is less than the median of the path priority list, the path is deemed insufficiently important to justify the cost of delayed decoding, and the processing flag is set to "discard directly." Otherwise, the path is set to "delayed retention" and enters the decoding queue.
[0073] This priority-driven discrimination mechanism ensures that system resources are more concentrated on key modal data on important paths, thereby improving global data utilization and key flow accuracy.
[0074] Secondly, for the modal data entries that have been set to "delay retention", that is, the data in the decoding queue, decoding and backfill operations need to be performed according to the original compression strategy and state path information:
[0075] First, the decoder corresponding to the compression strategy ID is called to decode the compressed format back to the original modal structure format. This process is consistent with the modal compression mechanism defined in step S1, ensuring that the decoded data structure can be seamlessly integrated into the main processing flow processing chain;
[0076] For each decoded modal data, the original path mark saved in the modal state transition diagram is further read as the backfill target path number to ensure that the data flow direction is consistent with the original design logic.
[0077] Subsequently, the modal data that has been decoded and marked with path numbers is sent to the main processing flow inlet channel in the order of path numbers, and a "delay cache" data source label is added to it to facilitate subsequent identification and scheduling optimization.
[0078] During actual operation, this backfill process allows the main processing flow to obtain as much effective information compensation as possible without being interfered with by low-priority data, thereby improving the overall processing accuracy.
[0079] As can be seen, through the collaborative processing of the above content, the present invention not only achieves latency-aware management of cached data, but also ensures the continuity and data integrity of the main processing flow. In particular, when modal data compression mechanisms introduce a large amount of latency risk, the proposed three-layer logic of "delay tolerance-priority identification-path backfilling" greatly enhances the system's overall scheduling resilience and resource utilization efficiency for heterogeneous modal flows, providing an optimization solution that balances efficiency and reliability for complex modal fusion systems.
[0080] Furthermore, this embodiment also provides a multimodal data perception and processing system for the sound intelligent membrane, such as Figure 2 Shown, including,
[0081] The modal diagram construction module constructs a modal state transition diagram based on the collected multimodal data and calibrates the cache nodes of abnormal data in the state transition path;
[0082] The spatial encoding module constructs a spatial position matrix based on the spatial position relationship of the sensor array within the sound-intelligence membrane structure. Using the partition numbers in the spatial position matrix and the strong interference identification threshold, it determines the interference areas in the audio and image modalities and labels the interference modalities.
[0083] The modal compression module switches the involved modal data to the compression processing state according to the preset modal compression strategy rules when the modal type of a partition in the interference modal label table is marked as an interference state, or the corresponding node in the modal state transition diagram has been marked as an abnormal cache node;
[0084] The delayed decoding module combines the compression state mark and the spatial position matrix index of the corresponding position to selectively decode or discard the modal data in the cache node according to the set delay, and backfills it to the main processing flow based on the state transfer path.
[0085] This embodiment also provides a computer device, which is applicable to the multimodal data perception and processing method of the sound-intelligence membrane, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the multimodal data perception and processing method for the sound-intelligence membrane proposed in the above embodiment.
[0086] The computer device may be a terminal, comprising a processor, memory, a communication interface, a display, and an input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and computer programs. The internal memory provides an environment for the operating system and computer programs stored in the non-volatile storage media. The communication interface of the computer device is used to communicate with external terminals via wired or wireless communication. Wireless communication may be achieved via Wi-Fi, a carrier network, NFC (near-field communication), or other technologies. The display of the computer device may be a liquid crystal display or an electronic ink display. The input device may be a touchscreen overlay on the display, buttons, a trackball, or a touchpad on the computer device housing, or an external keyboard, touchpad, or mouse.
[0087] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, it implements the multimodal data perception and processing method for the sound intelligent membrane proposed in the above embodiment.
[0088] In summary, by constructing a modal state transition diagram and a spatial position matrix, the present invention can accurately identify abnormal states and interference areas, and dynamically adjust the processing strategy, thereby effectively improving the system's response speed and processing flexibility to abnormal modal data. Especially in the case of frequent interference data or limited bandwidth resources, the present invention greatly reduces the data transmission and computing burden through modal compression and delayed decoding mechanisms, ensuring the continuity and stability of the main processing flow. At the same time, the method has strong adaptive capabilities, and can dynamically adjust cache management and path priority according to real-time state changes, effectively preventing data congestion and delay diffusion. In summary, the present invention significantly improves the resource utilization efficiency and robustness of the multimodal processing system while ensuring data perception accuracy.
[0089] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A multimodal data perception and processing method for a sound-intelligence membrane, characterized by: include: Construct a modal state transition diagram based on the collected multimodal data and identify cache nodes of abnormal data in the state transition path; Based on the spatial position relationship of the sensor array within the sound-intelligence membrane structure, a spatial position matrix is constructed. The interference areas in the audio and image modalities are determined and labeled using the partition numbers and strong interference identification thresholds in the spatial position matrix. When the modal type of a partition in the interference modal label table is marked as an interference state, or the corresponding node in the modal state transition diagram has been marked as an abnormal cache node, the modal data involved is switched to the compression processing state according to the preset modal compression strategy rules; Combining the compression state mark and the spatial position matrix index of the corresponding position, the modal data in the cache node is selectively decoded or discarded according to the set delay, and backfilled to the main processing flow according to the state transition path; The switching of the involved modal data to the compression processing state includes: screening out a set of cache node numbers that have been marked as abnormal states based on the partition number and modal type in each record in the interference modal label table, combined with the state encoding vector of the corresponding node in the modal state transition diagram; establishing a compression state switching mapping table based on the cache node number set; fusing the compression state switching mapping table with the modal state transition diagram, performing positioning and updating according to the node number, marking the compression state flag in the modal state transition diagram, and adding the corresponding compression strategy number to form an updated modal state transition diagram; sequentially reading the number of each abnormal node that has been set to the compression state, comparing it with the modal data stored in the cache node, selecting a corresponding compression structure according to the modal type, performing compression, and synchronously updating the cache node; The selective decoding or discarding of modal data in the cache node according to the set delay includes: searching the compression state switching table for a set of modal types currently in compression, extracting matching maximum tolerable delay parameters based on corresponding compression strategy IDs, and constructing a compression modality-delay threshold mapping table; locating the coordinate index of each cache node number in the spatial topology through the spatial matrix index based on the spatial position matrix in combination with the cache node number set, and obtaining the modal data set corresponding to the cache area; extracting the timestamp of each data record in the modal data set and comparing it with the current system time; if the timestamp is less than the current system time minus the maximum delay threshold corresponding to the compression modality-delay threshold mapping table, marking the corresponding data as expired and adding it to a pending cache list; extracting the path priority score of the corresponding node based on the updated modal state transition diagram for each cache record in the pending cache list, and comparing it with the median of the path priority list; when the priority score is less than the median of the path priority list, it is considered that the importance of the path is insufficient to support the delayed decoding cost, and the processing flag is set to directly discard; otherwise, it is set to delayed retention and enters the decoding queue.
2. The multimodal data perception and processing method for a sound-intelligence membrane according to claim 1, characterized in that: The calibration of the cache node includes: Extract the main feature parameter sequence representing the state evolution from the multimodal state; construct the state encoding vector by combining the main features of the multimodal state at the current moment. , and construct a state encoding sequence; Convert the state encoding sequence into a node state sequence, construct directed edges in the time direction, and form a modal state transition graph, where each node represents a state encoding vector and each edge represents a state transition under a timestamp. Abnormal state transition detection is performed in the modal state transition diagram path, and abnormal cache nodes are calibrated.
3. The multimodal data perception and processing method for a sound-intelligence membrane according to claim 2, characterized in that: The formula for abnormal state transition detection is as follows: , in, is the L1 norm, and represents the modal state vector at two consecutive time points, Indicates time arrive The change amplitude of the state encoding vector between , then the modal data corresponding to the current time point is cached as an abnormal candidate data block.
4. The multimodal data perception and processing method for a sound-intelligence membrane according to claim 1, characterized in that: The construction of the spatial position matrix includes: Based on the actual layout coordinates of each modal sensor in the sound-intelligence membrane structure, the geometric coordinate information of all sensors in the two-dimensional membrane surface is extracted to construct the initial coordinate set. , where the initial coordinate set Each element of consists of a sensor number and a position pair in the membrane surface coordinate system (X, Y); The initial coordinate set Map rows and columns in numerical order, divide the two-dimensional membrane surface into several spatial partitions, and based on the initial coordinate set in each partition Element aggregation generates a spatial position matrix , where the spatial position matrix Each element in contains the partition number and the number of modal sensors distributed in the corresponding block.
5. The multimodal data perception and processing method for a sound-intelligence membrane according to claim 4, characterized in that: The determination of the interference area in the audio and image modalities includes: In the spatial position matrix Extract all partition sets containing multi-source modal sensors, and construct a multi-modal interference candidate partition set by combining the node state indicators in the modal state transition graph; An interference strength index is determined for each partition in the multimodal interference candidate partition set, and an interference determination threshold is set to mark the partition whose interference strength index exceeds the interference determination threshold as an interference area.
6. The multimodal data perception and processing method for a sound-intelligence membrane according to claim 5, characterized in that: The process of determining the interference intensity index includes: Perform time domain feature analysis on the audio modal data segments within each spatial partition and extract the audio waveform disturbance level parameters; Perform inter-frame pixel fluctuation density analysis on the image modality data of each spatial partition to form image grade parameters; The audio waveform disturbance level parameter and the image level parameter are combined to form a disturbance intensity index.
7. The multimodal data perception and processing method for a sound-intelligence membrane according to claim 1, characterized in that: The backfilling to the main processing flow includes: For the modal data entries that have been set to be retained for delay, the decoder corresponding to the compression strategy ID is called to decode the compressed format back to the original modal structure format; For each decoded modal data, the original path mark saved in the modal state transition diagram is further read as the backfill target path number; The modal data that has been decoded and marked with path numbers is sent to the main processing flow inlet channel in the order of path numbers, and the delayed cache data source label is added at the same time.
8. A multimodal data perception and processing system for a sound-intelligence membrane, based on the multimodal data perception and processing method for a sound-intelligence membrane according to any one of claims 1 to 7, characterized in that: Also includes: The modal diagram construction module constructs a modal state transition diagram based on the collected multimodal data and calibrates the cache nodes of abnormal data in the state transition path; The spatial encoding module constructs a spatial position matrix based on the spatial position relationship of the sensor array within the sound-intelligence membrane structure. Using the partition numbers in the spatial position matrix and the strong interference identification threshold, it determines the interference areas in the audio and image modalities and labels the interference modalities. The modal compression module switches the involved modal data to the compression processing state according to the preset modal compression strategy rules when the modal type of a partition in the interference modal label table is marked as an interference state, or the corresponding node in the modal state transition diagram has been marked as an abnormal cache node; The delayed decoding module combines the compression state mark and the spatial position matrix index of the corresponding position to selectively decode or discard the modal data in the cache node according to the set delay, and backfills it to the main processing flow based on the state transfer path.
Citation Information
Patent Citations
Infrared sensor structure, display screen structure and terminal
CN109743419A
Environment sensing system based on intelligent driving
CN118850119A