Multi-modal data sensing processing method and system for sound intelligent film

By constructing a model state transition diagram and spatial position matrix, combining compression and delay decoding mechanisms, the information accuracy and stability problems of multimodal perception technology in the dynamic interference environment in the audio-intelligent membrane structure are solved, and the accurate identification and flexible processing of abnormal data are achieved, which improves the system's resource utilization efficiency and robustness.

CN120281754AActive Publication Date: 2025-07-08NANJING PIONE HIGH TECH

Patent Information

Application Number
CN202510764828.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing multimodal perception technology is difficult to maintain effective performance in the sound-intelligent membrane structure in an environment with frequent local interference, resulting in difficult to ensure information accuracy and stability. In particular, audio and image modalities are limited by physical propagation paths and structural occlusions, and abnormal data in local areas are prone to misjudgment, causing chain errors in the processing flow.

Method used

Construct a model state transition diagram and spatial position matrix, and calibrate the cache nodes and interference areas of exception data, combine compression strategies and delay decoding mechanisms to dynamically adjust the processing strategies to achieve accurate identification and processing of exception states and interference areas.

Benefits of technology

It improves the system's response speed and processing flexibility to abnormal mode data, reduces the burden of data transmission and calculation, ensures the continuity and stability of the main processing flow, and improves resource utilization efficiency and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281754A_ABST
    Figure CN120281754A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data sensing processing method and system for a sound intelligent film, and relates to the technical field of data processing. By constructing the modal state transition diagram and the spatial position matrix, the abnormal state and the interference area can be accurately identified, and the processing strategy is dynamically adjusted, so that the response speed and the processing flexibility of the system to the abnormal modal data are effectively improved. Particularly, under the condition that interference data is frequently transmitted or bandwidth resources are limited, data transmission and calculation burdens are greatly reduced through a modal compression and delay decoding mechanism, and continuity and stability of a main processing stream are guaranteed. Meanwhile, the method has high adaptive capacity, cache management and path priority can be dynamically adjusted according to real-time state changes, and data congestion and delay diffusion are effectively prevented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, in particular to a multi-modal data perception and processing method and system for a sound-intelligent membrane. Background Art

[0002] In the process of multi-modal data processing, how to integrate heterogeneous modal data such as images, audio, temperature and humidity, vibration, etc., to form a collaborative perception and dynamic response mechanism is one of the key issues in the design of current intelligent systems. In recent years, with the development of acoustic control technology and soft structural materials, the air film structure (i.e., sound-intelligent membrane) integrating sound function and intelligent perception ability has gradually become an important component in new type of expandable space, with advantages such as light weight, flexible deployment, and high functional integration.

[0003] Under this background, embedding multi-modal data processing strategies into the sound-intelligent membrane to achieve adaptive recognition and response to environmental states, personnel activities, and noise sources has become a research hotspot. However, since the sound-intelligent membrane structure is mostly in a semi-closed or highly reverberant environment, the perception signals are often accompanied by problems such as non-stationary interference, spatial interference overlap, and modal information drift, making it difficult for existing perception and processing methods to meet the dual requirements of accuracy and stability.

[0004] In existing multi-modal perception technologies, parallel modeling, attention mechanism fusion, or data-level alignment methods are often used to improve the adaptability and information integrity between different modal data. However, such methods often rely on stable signal input conditions and are difficult to maintain effective performance in an environment with frequent dynamic and local interferences. For example, in the sound-intelligent membrane structure, due to the limitations of the physical propagation path and structural occlusion of audio and image modalities, abnormal data in a local area is extremely likely to cause misjudgment, which in turn leads to chain errors in the processing flow. Summary of the Invention

[0005] In view of the problems existing in the above background art, the present invention is proposed.

[0006] Therefore, the problem to be solved by the present invention is how to construct a multi-modal perception framework that can transfer with the time state, support abnormal node marking, and has the ability of position association and compression selection processing on the premise of ensuring information integrity.

[0007] To solve the above technical problems, the present invention provides the following technical solutions: In the first aspect, the present invention provides a multimodal data perception and processing method for a sound intelligence membrane, which includes constructing a modal state transition diagram based on the collected multimodal data, and calibrating the cache nodes of abnormal data in the state transition path; constructing a spatial position matrix based on the spatial position relationship of the sensor array in the sound intelligence membrane structure, and determining the interference area in the audio and image modes and marking the interference modal label through the partition number and strong interference identification threshold in the spatial position matrix; when the modal type of a partition in the interference modal label table is marked as an interference state, or the corresponding node in the modal state transition diagram has been marked as an abnormal cache node, the modal data involved is switched to a compressed processing state according to a preset modal compression strategy rule; combined with the compression state mark and the spatial position matrix index of the corresponding position, the modal data in the cache node is selectively decoded or discarded according to the set delay, and backfilled to the main processing flow according to the state transition path.

[0008] As a preferred solution of the multimodal data perception processing method for the sound intelligence membrane described in the present invention, the calibration of the cache node includes: extracting the main feature parameter sequence representing the state evolution from the multimodal; constructing the state encoding vector with the main feature combination of the multimodal at the current moment , and construct a state coding sequence; convert the state coding sequence into a node state sequence, construct directed edges in the time direction, and form a modal state transition diagram, in which each node represents a state coding vector and each edge represents a state transition under a timestamp; perform abnormal state transition detection in the modal state transition diagram path and calibrate abnormal cache nodes.

[0009] As a preferred solution of the multimodal data perception and processing method for the sound intelligence membrane of the present invention, the formula for abnormal state transition detection is as follows: , in, is the L1 norm, and represents the modal state vector at two consecutive time points, Indicates at time arrive The change amplitude of the state encoding vector between , then the modal data corresponding to the current time point is cached as an abnormal candidate data block.

[0010] As a preferred embodiment of the multimodal data perception and processing method for the sound intelligence membrane of the present invention, the construction of the spatial position matrix includes: based on the actual layout coordinate position of each modal sensor in the sound intelligence membrane structure, extracting the geometric coordinate information of all sensors in the two-dimensional membrane surface, and constructing an initial coordinate set , where the initial coordinate set Each element of consists of a sensor number and a position pair in the membrane surface coordinate system (X, Y); the set of the initial coordinates is subjected to row-column mapping in the order of the numbers, the two-dimensional membrane surface is divided into several spatial partitions, and based on the set of the initial coordinates within each partition the aggregation condition of the elements generates a spatial position matrix , wherein, each element in the spatial position matrix contains a partition number and the number of modal sensor distributions in the corresponding block.

[0011] As a preferred solution of the multi-modal data perception and processing method for the sound-intelligent membrane according to the present invention, wherein: the interference regions in the determined audio and image modalities include: in the spatial position matrix extract all the partition sets containing multi-source modal sensors, combine the node state indexes in the modal state transition diagram, and construct a multi-modal interference candidate partition set; determine an interference intensity index for each partition in the multi-modal interference candidate partition set, and by setting an interference determination threshold, mark the partitions with the interference intensity index exceeding the interference determination threshold as interference regions.

[0012] As a preferred solution of the multi-modal data perception and processing method for the sound-intelligent membrane according to the present invention, wherein: the determination process of the interference intensity index includes: performing time-domain feature analysis on the audio modal data segments within each spatial partition, and extracting an audio waveform perturbation level parameter; performing an analysis on the frame-to-frame pixel fluctuation density of the image modal data of each spatial partition to form an image level parameter; combining the audio waveform perturbation level parameter and the image level parameter to form an interference intensity index.

[0013] As a preferred solution of the multi-modal data perception and processing method for the sound-intelligent membrane according to the present invention, wherein: the switching of the involved modal data to the compression processing state includes: according to the partition numbers and modal types in each record in the interference modal label table, combining the state coding vectors of the corresponding nodes in the modal state transition diagram, screening out the set of cache node numbers that have been calibrated as abnormal states; based on the set of cache node numbers, establishing a compression state switching mapping table; performing a fusion process on the compression state switching mapping table and the modal state transition diagram, performing positioning update according to the node numbers, marking a compression state flag bit in the modal state transition diagram, and attaching the corresponding compression strategy number to form an updated modal state transition diagram; sequentially reading each abnormal node number that has been set to the compression state, comparing with the modal data stored in the cache node, selecting a corresponding compression structure according to the modal type, performing compression, and synchronously updating the cache node.

[0014] As a preferred embodiment of the multimodal data perception and processing method for the sound-intelligent film of the present invention, wherein: the selective decoding or discarding of the modal data in the cache node according to a set time delay includes: retrieving the set of modal types currently in the compressed state in the compression state transition table, and extracting the matching maximum tolerable time delay parameter according to the corresponding compression policy ID to construct a compressed mode-time delay threshold mapping table; according to the spatial position matrix, combining with the set of cache node numbers, for each cache node number, locating the coordinate index in the spatial topology through matrix indexing to obtain the set of modal data in the corresponding cache area; for each data record in the set of modal data, extracting the timestamp and comparing it with the current system time. If the timestamp is less than the current system time minus the maximum time delay threshold corresponding to the compressed mode-time delay threshold mapping table, the corresponding data is marked as expired and added to the cache list to be processed; for each cache record in the cache list to be processed, according to the updated modal state transition graph, extracting the path priority score of the corresponding node and comparing it with the median of the path priority list: when the priority score is less than the median of the path priority list, it is considered that the importance of this path is not sufficient to support the delay decoding cost, and the processing flag is set to directly discard; otherwise, it is set to delay retention and enters the decoding queue.

[0015] As a preferred embodiment of the multimodal data perception and processing method for the sound-intelligent film of the present invention, wherein: the backfilling to the main processing stream includes: for the modal data entries that have been set to delay retention, calling the decoder corresponding to the compression policy ID to decode from the compressed format back to the original modal structure format; for each decoded modal data, further reading the original path label saved in the modal state transition graph as the backfilling target path number; sending the modal data that has been decoded and labeled with the path number into the main processing stream entry channel in the order of the path number, and adding a delay cache data source label at the same time.

[0016] In a second aspect, the present invention provides a multimodal data perception and processing system for a sound-intelligent film, which includes: The modality graph construction module constructs a modality state transition graph based on the collected multi-modal data, and calibrates the cache nodes of abnormal data in the state transition path; the spatial encoding module constructs a spatial position matrix based on the spatial position relationship of the sensor array in the sound-intelligence membrane structure, and determines the interference regions in the audio and image modalities and labels the interference modality tags through the partition numbers and strong interference identification thresholds in the spatial position matrix; the modality compression module, when the modality type of a certain partition in the interference modality tag table is marked as an interference state, or the corresponding node in the modality state transition graph has been calibrated as an abnormal cache node, switches the involved modality data to the compression processing state according to the preset modality compression strategy rules; the delay decoding module, in combination with the compression state mark and the spatial position matrix index of the corresponding position, selectively decodes or discards the modality data in the cache node at a set time delay, and backfills it to the main processing stream according to the state transition path.

[0017] The beneficial effects of the present invention are as follows: By constructing a modality state transition graph and a spatial position matrix, the present invention can accurately identify abnormal states and interference regions, and dynamically adjust the processing strategy, thereby effectively improving the response speed and processing flexibility of the system to abnormal modality data. Especially in the case of frequent interference data or limited bandwidth resources, the present invention significantly reduces the data transmission and calculation burden through the modality compression and delay decoding mechanisms, ensuring the continuity and stability of the main processing stream. At the same time, the method has strong adaptability, can dynamically adjust the cache management and path priority according to real-time state changes, and effectively prevent data congestion and delay diffusion. In summary, the present invention significantly improves the resource utilization efficiency and robustness of the multi-modal processing system while ensuring the data perception accuracy. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 It is a flowchart of a multi-modal data perception processing method for a sound-intelligence membrane.

[0020] Figure 2 It is a structural diagram of a multi-modal data perception processing system for a sound-intelligence membrane. Detailed Embodiments

[0021] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention in conjunction with the drawings in the specification.

[0022] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0023] Secondly, as used herein, an "embodiment" or "embodiments" refers to specific features, structures, or characteristics that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that is mutually exclusive with other embodiments.

[0024] As mentioned in the above background art, embedding multi-modal data processing strategies into the sound-intelligent membrane to achieve adaptive recognition and response to environmental states, personnel activities, and noise sources has become a research hotspot. However, due to the fact that the sound-intelligent membrane structure is mostly in a semi-closed or highly reverberant environment, the perceived signals are often accompanied by problems such as non-stationary interference, spatial interference overlap, and modal information drift, making it difficult for existing perception processing methods to meet the dual requirements of accuracy and stability. In existing multi-modal perception technologies, parallel modeling, attention mechanism fusion, or data-level alignment methods are often used to improve the adaptability and information integrity between different modal data. However, such methods often rely on stable signal input conditions and are difficult to maintain effective performance in an environment with frequent dynamic and local interferences. For example, in the sound-intelligent membrane structure, due to the limitations of the physical propagation path and structural occlusion of the audio and image modalities, abnormal data in a local area is extremely likely to cause misjudgment, thereby triggering a chain of errors in the processing flow.

[0025] Figure 1 It is a flowchart of a multi-modal data perception processing method for a sound-intelligent membrane according to an embodiment of the present invention. As Figure 1 shown, in the multi-modal data perception processing method for a sound-intelligent membrane, it includes: S1: Construct a modal state transition graph based on the collected multi-modal data, and mark the cache nodes of abnormal data in the state transition path.

[0026] S1.1: Extract the main feature parameter sequences representing state evolution from the three modalities of audio, image, and vibration; construct a state coding vector with the main feature combination of the three modalities at the current moment , where each component is encoded into a finite state set (such as "low, medium, high" corresponding to 0 / 1 / 2) through a hierarchical mapping method (such as thresholds, intervals), so that the state coding vector can be used to construct a state coding sequence.

[0027] S1.2: The state coding sequence Convert it into a node state sequence, construct directed edges in the time direction to form a modal state transition graph, where each node represents a state encoding vector and each edge represents a state transition at a time stamp.

[0028] Among them, only the occurrence frequency of each transition path is retained in the modal state transition graph. As a basic attribute. Based on the path frequency, construct a path priority list, where the higher the frequency, the higher the priority. This priority will be used as the basis for the path selection strategy in the interference state backfill.

[0029] S1.3: In the modal state transition graph path, use the following abnormal state transition detection formula to calibrate abnormal cache nodes: , Among them, is the L1 norm, and represent the modal state vectors at two consecutive time points, represents at time to the change amplitude of the state encoding vector between; if (the threshold is determined by the historical average transition level of the training set), then cache the modal data corresponding to the current time point as an abnormal candidate data block, and record the node number and the occurring mode.

[0030] S2: Based on the spatial position relationship of the sensor arrays within the sound-intelligence membrane structure, construct a spatial position matrix, and determine the interference regions in the audio and image modalities and label the interference mode labels through the partition numbers and strong interference identification thresholds in the spatial position matrix.

[0031] S2.1: Construct a spatial position matrix.

[0032] Based on the actual layout coordinate positions of the various modal sensors in the sound-intelligence membrane structure, extract the geometric coordinate information of all sensors in the two-dimensional membrane surface to construct an initial coordinate set , where each element of consists of a sensor number and a position pair in the membrane surface coordinate system (X, Y).

[0033] The initial coordinate set Perform row-column mapping in sequential order, and divide the two-dimensional membrane surface into several spatial partitions through spatial clustering and region coding algorithms. There are various ways to divide the spatial region in operation. Among them, the Voronoi diagram method is suitable for processing the adaptive division of partition boundaries under non-uniform distribution, and can make each partition enclose the nearest neighbor sensor nodes as much as possible; while DBSCAN can effectively generate distribution regions in the scenario of uneven density, which is beneficial to dynamically identify abnormal cluster distribution regions. The method to be adopted is not uniquely limited, and can be flexibly selected according to the actual deployment situation.

[0034] And based on the initial coordinate set within each partition Generate a spatial position matrix based on the element aggregation situation , where each element in the spatial position matrix contains the partition number and the number of modal sensor distributions in the corresponding block.

[0035] Furthermore, among the abnormal nodes calibrated in the state transition diagram, if their positions in the physical space cannot be determined, it is impossible to realize interference recognition based on the spatial aggregation effect. Therefore, the present invention introduces a spatial partition-modal state joint coding field in the data structure of each abnormal node. This coding field includes three sub-fields: the number of the spatial block to which it belongs, the sensor modal type, and the corresponding state coding.

[0036] S2.2: Determine the interference regions in the audio and image modalities and label the interference modal labels.

[0037] After generating the spatial position matrix containing the modal type distribution, the system starts to screen the interference candidate regions. The key here is to identify "multi-modal coupling interference", that is, the situation where there are abnormal nodes in multiple modalities in the same region.

[0038] During the operation, first extract all the partition sets containing multi-source modal sensors in the spatial position matrix , and combine the node state indicators in the modal state transition diagram to construct a multi-modal interference candidate partition set, where each partition is an aggregation area that simultaneously contains abnormal state nodes in terms of spatial position.

[0039] Determine the interference intensity index for each partition in the multi-modal interference candidate partition set. Among them, the interference intensity index is jointly perturbed according to the waveform instability index in the audio modality and the pixel region noise dispersion degree in the image modality. Mark the partitions with the interference intensity index exceeding the medium interference region as interference regions.

[0040] Label the modal types in all interference regions to the interference modal label table. Among them, each record in the interference modal label table contains the interference spatial partition number, the modal type, and the interference intensity index, and write back this label table to the node attributes of the original modal state transition diagram.

[0041] Preferably, the generation process of the interference intensity index includes: Perform time-domain feature analysis on the audio modal data segments in each spatial partition to extract the audio waveform perturbation level parameters.

[0042] Specifically, first divide the audio signal into multiple sub-segments according to a preset time window, and execute a stability determination process for each sub-segment: If the continuous fluctuation amplitude of the signal energy in any sub-segment exceeds twice the average value of the entire time window, and this abnormal fluctuation continues for more than three sub-segments, then this audio segment is marked as a strong perturbation segment; If there is only one abnormal fluctuation and it lasts for less than two sub-segments, it is marked as a light perturbation segment; Otherwise, it is marked as a medium perturbation segment. After counting the proportion of each type of perturbation segment, according to whether the proportion of the strong perturbation segment exceeds U% of the total number of segments, where the U value can determine the optimal threshold through the ROC curve to distinguish the interference and non-interference regions. Further classification is carried out: If it exceeds, the audio perturbation level of this partition is determined to be the audio high-perturbation area; If the sum of the strong perturbation segment and the medium perturbation segment exceeds 2U%, it is marked as the audio medium-perturbation area; Otherwise, it is the audio low-perturbation area.

[0043] Perform frame-by-frame pixel fluctuation density analysis on the image modal data of each spatial partition to form an image level parameter.

[0044] First, perform region matching and differential calculation on consecutive frame images to extract the total area ratio of the pixel gray-scale variation regions of the target region between any three consecutive frames. If this ratio exceeds 20% of the entire image area three times in a row (this data is only preset data and should be based on actual operations in applications), it is determined as a high-fluctuation segment; If it is not continuous but the occurrence frequency exceeds once, it is a medium-fluctuation segment; Otherwise, it is a low-fluctuation segment. Based on this, count the distribution density and spatial concentration degree of each type of fluctuation segment. If the high-fluctuation segment continuously appears in the center of the image area for more than five frames and covers more than h% of the center area (this data should be based on actual operations in applications), then this image area is marked as the image strong-perturbation area; If the high-fluctuation segment is mainly concentrated on the edge or only appears briefly (for example, the number of frames does not exceed three frames), it is marked as the image light-perturbation area; Other situations are classified as the image medium-perturbation area.

[0045] After obtaining the audio disturbance level and the image disturbance level, combine the two and generate the interference intensity index label for the current spatial partition according to the following rules: If the audio disturbance level is a high-disturbance area and the image disturbance level is a strong-disturbance area or a medium-disturbance area, the current partition is marked as a strong interference area; if the audio level is a medium-disturbance area and the image level is a strong-disturbance area or a medium-disturbance area, it is marked as a medium interference area; if the audio level is a low-disturbance area but the image level is a strong-disturbance area, it is still marked as a medium interference area; only when both are low-disturbance areas, it is marked as a non-interference area. In addition, if the image disturbance level is a strong-disturbance area, but its fluctuation area is always limited to the edge area and does not affect the recognition quality of the central part of the image, the original strong interference label is downgraded to a medium interference.

[0046] It should be noted that the vibration mode characteristics are mainly used for state sequence modeling and path branch identification and do not participate in the determination of the disturbance intensity index.

[0047] In view of the high coincidence of the input data in the process of constructing the modal state transition diagram and the process of generating the spatial matrix, to improve the processing efficiency, the operations in S1 and S2 can be executed simultaneously through a thread-level parallel mechanism. However, since both involve read and write operations on node states and structure identifiers, to avoid data competition and state conflicts, the present invention introduces a lightweight distributed lock mechanism for synchronization. The specific mechanism is as follows: Each node state record structure is appended with a state lock flag, and the S1 and S2 sub-threads apply for access permissions before read and write operations; the state write operation adopts a write priority strategy, and the read operation adopts a locked snapshot mechanism to avoid inconsistent propagation; a distributed conflict mediation strategy based on timestamps is adopted to ensure that concurrent tasks are consistent at both the spatial topology and state flow conversion levels.

[0048] S3: When the modal type of a partition in the interference modal label table is marked as an interference state, or the corresponding node in the modal state transition diagram has been calibrated as an abnormal cache node, switch the involved modal data to the compression processing state according to the preset modal compression strategy rules.

[0049] When it is detected that any record in the interference modal label table indicates that there is an interference anomaly in a certain block in terms of audio, image, vibration or other modalities, or a certain node in the modal state transition diagram has shown a persistent abnormal cache state, it indicates that the modal data can no longer accurately express information according to the conventional path or has no redundant interference filtering ability. On this basis, the present invention designs a dual-trigger mechanism based on "interference trigger + cache anomaly" to ensure a high-timeliness response to data anomalies in a multi-source modal system. Different from traditional systems that often regard modal anomalies as a post-processing stage, the present invention moves the compression processing forward to the data stream distribution center stage, effectively avoiding modal pollution and cache resource waste caused by anomaly propagation.

[0050] To implement the above switching process, first, according to the partition numbers and modal types in each record of the interference modal label table, combined with the state coding vectors of the corresponding nodes in the modal state transition diagram, all the cache node number sets that have been calibrated as abnormal states are screened out. This operation is essentially a cross-fusion of the calibration results of data from two sources. On the one hand, it comes from the historical state evolution chain in the modal state transition diagram, and on the other hand, it comes from the interference label table jointly determined by the spatial position matrix and the modal recognition model in the previous stage. The combination of the two can form a highly credible set of compressed candidate nodes. During the operation process, a fast screening path is established through a hash mapping function, so that the compression recognition time of nodes in the large-scale modal graph does not exceed linear complexity, enhancing the processing efficiency.

[0051] Based on the cache node number set, a compressed state switching mapping table is established. Each mapping record in it contains: abnormal node number, modal type, the number of the spatial partition it belongs to, and the number of the compression strategy to be adopted. It is worth emphasizing that this mapping table is not a static structure, but is generated in real time by combining dynamic variables such as the current operating load status, adaptive network bandwidth, and cache availability rate to ensure an appropriate match of the compression strategy. For example, in a high-concurrency interference state, the compression strategy number will tend to a high compression ratio and high-robustness structure, while in a normal background, a lightweight strategy is preferred first to maintain information fidelity.

[0052] Among them, the selection of the compression strategy number is based on the preset modal compression strategy rules, and corresponding compression methods are respectively matched according to the modal types (audio, image, vibration). For the audio modality, methods such as low-frequency information discard, wavelet packet feature coding, and speech main contour extraction are given priority. The core is to weaken the interference of background noise and retain the main form of the signal. For the image modality, compression methods such as regional pixel aggregation, edge sparse reconstruction, or low-resolution version replacement are adopted, aiming to reduce the visual pollution caused by pixel interference. For the vibration modality, strategies such as main frequency band focusing, redundant cycle removal, and Fourier peak clipping are adopted to compress the data volume from the frequency spectrum dimension. All strategies are equipped with compression ratio selection, allowing the compression level to be automatically selected according to the current network load. For example, the product of the predicted transmission delay function and the cache read frequency is used as a dynamic reference for the compression factor.

[0053] Subsequently, the compressed state transition mapping table is fused with the modal state transition graph, located and updated according to the node numbers. The compressed state flag bits are marked in the modal state transition graph, and the corresponding compression policy numbers are attached, so that the abnormal nodes in the modal state transition graph have compression attribute fields, forming an updated modal state transition graph, which is convenient for subsequent cache scheduling and data reading modules to distinguish and process. The technical key of this operation is to maintain the topological integrity of the original state transition graph while introducing compression flags, so that the subsequent scheduler can clearly distinguish between regular nodes and compressed nodes when reading the graph. Different from traditional systems that use additional compression channels for bypass processing, the present invention realizes unified management of data paths through in-graph structure embedded attribute updates, avoiding additional index overhead and concurrent consistency problems.

[0054] Furthermore, according to the records in the compressed state transition mapping table, the abnormal node numbers that have been set to the compressed state are read sequentially. Comparing with the modal data saved in the abnormal cache nodes, select the corresponding compression structure according to the modal type (such as audio modal compressed into a feature envelope, image modal compressed into a regional pixel encoding, vibration modal compressed into a main frequency parameter set) for compression, and synchronously update the cache nodes, marking the compression structure type and the compression version number.

[0055] After the compression structure replacement is completed, the compressed state update is pushed to the modal scheduler. Based on this push information, the scheduler re-plans the node scheduling path. In the subsequent data reading and cache scheduling phases, the compressed nodes are placed in the medium and low priority levels to avoid frequent calls to the compressed nodes in the case of limited resources, thereby reducing the impact of redundant data on system performance.

[0056] During subsequent operation, if it is detected that the interference intensity has dropped below the back-off threshold, the system automatically calls the preset compressed state recovery mechanism to switch the node state back to the normal data processing path and restore the original cache structure and scheduling logic.

[0057] S4: Combine the compressed state mark and the spatial position matrix index at the corresponding position to selectively decode or discard the modal data in the cache node according to the set time delay, and backfill it to the main processing stream according to the state transition path.

[0058] After completing the main processing stream and cache separation strategy in step S3, to achieve effective management and recycling of cache data, it is necessary to evaluate the time delay tolerance and determine the retention or deletion of the modal data in the buffered state based on the compressed state mark and the spatial position matrix index.

[0059] First, retrieve the set of modal types currently in the compressed state in the compressed state transition table, and extract the maximum tolerable time delay parameter matching it according to the compression policy ID corresponding to each modal entry in the set, thereby constructing a compressed modal - time delay threshold mapping table.

[0060] Next, based on the spatial position matrix constructed in step S2 and combined with the cache node number set, for each cache node number, its coordinate index in the spatial topology is located through matrix indexing, and then the modal data set of its corresponding cache area is obtained.

[0061] For each data record in the modal data set, its timestamp is extracted and compared with the current system time. If the timestamp is less than the maximum delay threshold corresponding to the compression mode - delay threshold mapping table of the current system time, that is, the data of the mode has exceeded the maximum delay threshold tolerated by the compression policy, then the data is marked as "expired" and added to the list of caches to be processed.

[0062] This marking process ensures that subsequent retention or deletion operations are only performed on the actually timed-out modal data, avoiding unnecessary processing overhead for valid data.

[0063] Subsequently, for each cache record in the list of caches to be processed, according to the updated modal state transition graph constructed in step S3, the path priority score of the corresponding node is extracted and compared with the median of the path priority list. When the priority score is less than the median of the path priority list, it is considered that the importance of this path is not sufficient to support the delay decoding cost, so the processing flag is set to "directly discard"; otherwise, it is set to "delay retention" and enters the decoding queue.

[0064] This priority-driven discrimination mechanism ensures that system resources are more concentrated on the key modal data on important paths, thereby improving the global data utilization rate and the accuracy of key streams.

[0065] Secondly, for the modal data entries that have been set to "delay retention", that is, the data in the decoding queue, decoding and backfilling operations need to be performed according to the original compression policy and state path information:

[0066] First, call the decoder corresponding to the compression policy ID to decode from the compressed format back to the original modal structure format. This process is consistent with the modal compression mechanism defined in step S1 to ensure that the decoded data structure can be seamlessly connected to the main processing flow processing chain;

[0067] For each decoded modal data, further read the original path label saved in the modal state transition graph as the backfill target path number to ensure that the data flow direction is consistent with the original design logic.

[0068] Subsequently, the modal data that has been decoded and has the path number marked is sent into the main processing flow entry channel in the order of the path numbers, and at the same time, a "delayed cache" data source label is added to it for subsequent identification and scheduling optimization.

[0069] During the actual operation process, this backfilling process allows the main processing stream to obtain as much effective information compensation as possible without being interfered by low-priority data, thereby improving the overall processing accuracy.

[0070] It can be seen that through the collaborative processing of the above content, the present invention not only realizes the time-delay aware management of cached data, but also ensures the continuity and data integrity of the main processing stream. Especially in the case where the modal data compression mechanism introduces a large amount of latency risks, the proposed three-layer logic of "delay tolerance - priority identification - path backfilling" greatly enhances the overall scheduling resilience and resource utilization efficiency of the system for heterogeneous modal streams, providing an optimized solution that takes into account both efficiency and reliability for complex modal fusion systems.

[0071] Furthermore, the present embodiment also provides a multi-modal data perception processing system for the audio-intelligent film, as Figure 2 shown, including, A modal graph construction module that constructs a modal state transition graph based on the collected multi-modal data and calibrates the cache nodes of abnormal data in the state transition path; A spatial encoding module that constructs a spatial position matrix based on the spatial position relationship of the sensor array within the audio-intelligent film structure, and determines the interference regions in the audio and image modalities and labels the interference modal tags through the partition numbers and strong interference identification thresholds in the spatial position matrix; A modal compression module that, when the modal type of a certain partition in the interference modal tag table is marked as an interference state, or the corresponding node in the modal state transition graph has been calibrated as an abnormal cache node, switches the involved modal data to the compression processing state according to the preset modal compression strategy rules; A delay decoding module that, in combination with the compression state mark and the spatial position matrix index at the corresponding position, selectively decodes or discards the modal data in the cache node according to the set time delay, and backfills it to the main processing stream according to the state transition path.

[0072] The present embodiment also provides a computer device applicable to the case of the multi-modal data perception processing method for the audio-intelligent film, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the multi-modal data perception processing method for the audio-intelligent film proposed in the above embodiment.

[0073] The computer device may be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or may also be a button, a trackball, or a touchpad provided on the housing of the computer device, or may also be an external keyboard, a touchpad, or a mouse, etc.

[0074] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the multi-modal data perception and processing method for the sound intelligence film as proposed in the above embodiment.

[0075] In summary, by constructing a modal state transition diagram and a spatial position matrix, the present invention can accurately identify abnormal states and interference regions, and dynamically adjust the processing strategy, thereby effectively improving the response speed and processing flexibility of the system to abnormal modal data. Especially in the case of frequent interference data or limited bandwidth resources, the present invention significantly reduces the data transmission and calculation burden through a modal compression and delayed decoding mechanism, ensuring the continuity and stability of the main processing flow. At the same time, the method has strong adaptability and can dynamically adjust cache management and path priorities according to real-time state changes, effectively preventing data congestion and delay diffusion. In summary, while ensuring the data perception accuracy, the present invention significantly improves the resource utilization efficiency and robustness of the multi-modal processing system.

[0076] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A multi-modal data perception and processing method for an intelligent sound film, characterized in that: Including: Construct a modal state transition graph based on the collected multimodal data, and calibrate the cache nodes of abnormal data in the state transition path; Based on the spatial position relationship of the sensor array in the audio-intelligent membrane structure, construct a spatial position matrix, and determine the interference regions in the audio and image modalities and label the interference modality tags through the partition numbers and strong interference identification thresholds in the spatial position matrix; When the modality type of a certain partition in the interference modality tag table is marked as an interference state, or the corresponding node in the modal state transition graph has been calibrated as an abnormal cache node, switch the involved modal data to the compression processing state according to the preset modal compression strategy rules; Combining the compression state mark and the spatial position matrix index of the corresponding position, selectively decode or discard the modal data in the cache node according to the set time delay, and backfill it to the main processing stream according to the state transition path.

2. The multimodal data perception and processing method for the sound intelligence film according to claim 1, wherein: The calibration of the cache node includes: Extract the main feature parameter sequence representing the state evolution from the multi-modalities; construct the state encoding vector with the main feature combination of the multi-modalities at the current moment , and construct the state encoding sequence; Convert the state coding sequence into a node state sequence, construct a directed edge in the time direction to form a modal state transition graph, where each node represents a state coding vector and each edge represents a state transition under the time stamp; Perform abnormal state transition detection in the path of the modal state transition graph and calibrate the abnormal cache nodes.

3. The multimodal data perception and processing method for the sound-intelligent film according to claim 2, characterized in that: The formula for the abnormal state transition detection is as follows: , Among them, is the L1 norm, and represent the modal state vectors at two consecutive time points, represents the change amplitude of the state encoding vector from time to ; if , the modal data corresponding to the current time point is cached as an abnormal candidate data block.

4. The multimodal data perception and processing method for the sound-intelligent film according to claim 1, wherein: The construction of the spatial position matrix includes: Based on the actual layout coordinate positions of each modal sensor in the sound-intelligent membrane structure, extract the geometric coordinate information of all sensors within the two-dimensional membrane surface, and construct an initial coordinate set , where each element of the initial coordinate set consists of a sensor number and a position pair in the membrane surface coordinate system (X, Y); The set of initial coordinates is mapped row by row and column by column in the order of the numbers, dividing the two-dimensional membrane surface into several spatial partitions, and based on the set of initial coordinates in each partition, a spatial position matrix is generated according to the element aggregation situation. Among them, each element in the spatial position matrix contains the partition number and the number of modal sensor distributions in the corresponding block.

5. The multimodal data perception and processing method for the sound intelligence film according to claim 4, characterized in that: The determination of the interference regions in the audio and image modalities includes: Extract all partition sets containing multi-source modal sensors from the spatial position matrix and construct a multi-modal interference candidate partition set by combining the node state indicators in the modal state transition diagram. Determine the interference intensity index for each partition in the multimodal interference candidate partition set, and mark the partitions with the interference intensity index exceeding the interference determination threshold as interference regions by setting the interference determination threshold.

6. The multimodal data perception and processing method for the sound-intelligent film according to claim 5, wherein: The determination process of the interference intensity index includes: Perform time-domain feature analysis on the audio modal data segments in each spatial partition, and extract the audio waveform perturbation level parameters; Analyze the density of frame-to-frame pixel fluctuations of the image modal data in each spatial partition to form image level parameters; Combine the audio waveform perturbation level parameters and the image level parameters to form the interference intensity index.

7. The multimodal data perception and processing method for the sound-intelligent film according to claim 1, wherein: The switching of the involved modal data to the compression processing state includes: According to the partition numbers and modality types in each record in the interference modality tag table, combined with the state coding vectors of the corresponding nodes in the modal state transition graph, screen out the set of cache node numbers that have been calibrated as abnormal states; Based on the set of cache node numbers, establish a compression state switching mapping table; Fuse the compression state switching mapping table with the modal state transition graph, perform positioning and updating according to the node numbers, mark the compression state flag bit in the modal state transition graph, and attach the corresponding compression strategy number to form an updated modal state transition graph; Read the abnormal node numbers that have been set to the compression state in sequence, compare with the modal data stored in the cache node, select the corresponding compression structure according to the modality type, perform compression, and synchronously update the cache node.

8. The multimodal data perception and processing method for the sound-intelligent film according to claim 1, wherein: The selective decoding or discarding of the modal data in the cache node according to the set time delay includes: Retrieve the set of modality types currently in the compression state in the compression state switching table, and extract the matching maximum tolerable time delay parameters according to the corresponding compression strategy ID to construct a compression modality - time delay threshold mapping table; According to the spatial position matrix, combined with the set of cache node numbers, for each cache node number, locate the coordinate index in the spatial topology through the spatial matrix index, and obtain the set of modal data for the corresponding cache area; For each data record in the set of modal data, extract the timestamp and compare it with the current system time. If the timestamp is less than the maximum delay threshold corresponding to the compression mode - delay threshold mapping table of the current system time minus the compression mode, the corresponding data is marked as expired and added to the cache list to be processed; For each cache record in the cache list to be processed, according to the updated mode state transition diagram, extract the path priority score of the corresponding node and compare it with the median of the path priority list: When the priority score is less than the median of the path priority list, it is considered that the importance of this path is not sufficient to support the cost of delayed decoding, and the processing flag is set to be directly discarded; otherwise, it is set to be delayed and retained and enters the decoding queue.

9. The multimodal data perception and processing method for the sound-intelligent film according to claim 8, characterized in that: The backfilling to the main processing stream includes: For the modal data entries that have been set to be delayed and retained, call the decoder corresponding to the compression policy ID to decode from the compressed format back to the original modal structure format; For each decoded modal data, further read the original path label saved in the modal state transition diagram as the backfilling target path number; Send the modal data that has been decoded and has the path number marked in order of the path number into the main processing stream entry channel, and at the same time add the delayed cache data source label.

10. A multimodal data perception and processing system for a sound-intelligent film, based on the multimodal data perception and processing method for a sound-intelligent film according to any one of claims 1 to 9, characterized in that: It also includes: A modal graph construction module that constructs a modal state transition diagram based on the collected multi-modal data and marks the cache nodes of abnormal data in the state transition path; A spatial encoding module that constructs a spatial position matrix based on the spatial position relationship of the sensor array in the sound intelligence film structure, and determines the interference area in the audio and image modalities and marks the interference modality label through the partition number and strong interference identification threshold in the spatial position matrix; A modal compression module that, when a certain partition's modal type in the interference modality label table is marked as an interference state, or the corresponding node in the modal state transition diagram has been marked as an abnormal cache node, switches the involved modal data to the compression processing state according to the preset modal compression policy rules; A delay decoding module that, in combination with the compression state mark and the spatial position matrix index of the corresponding position, selectively decodes or discards the modal data in the cache node according to the set delay and backfills it to the main processing stream according to the state transition path.

Citation Information

Patent Citations

  • Infrared sensor structure, display screen structure and terminal

    CN109743419A

  • Environment sensing system based on intelligent driving

    CN118850119A

  • Data governance risk early warning method based on big data mining

    CN120066862A

Cited By

  • Adaptive filtering algorithm and device for multi-noise-source environment of sound intelligent film

    CN120496553A

  • Adaptive filtering algorithm and device for sound-intelligence membrane in multi-noise source environment

    CN120496553B

  • Smart park comprehensive data management method and system

    CN120951313A

  • A Smart Park Integrated Data Management Method and System

    CN120951313B

  • Industrial component defect area positioning method based on microscopic image transfer learning

    CN121639663A