A city soundscape adaptive regulation method and system based on multi-source data
By processing multi-source data and calculating sound field delay phasors, a dynamic virtual cluster topology network is generated, which solves the problem of fragmented defense surface for sudden sound energy leakage events in urban soundscape control, and realizes precise suppression of sound energy and intelligent upgrading of urban acoustic governance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XUCHANG LINQIYU ECOLOGICAL GARDEN CO LTD
- Filing Date
- 2026-03-16
- Publication Date
- 2026-06-02
AI Technical Summary
In urban soundscape control, when faced with sudden sound energy leakage events in open square areas, the lack of effective network joint orchestration array data reshaping logic and joint network control response leads to fragmented failure of the noise defense surface, making it impossible to achieve tight network stabilization with full-dimensional network coverage.
By acquiring acoustic audio and visual image data synchronously from sensing units deployed at urban spatial nodes, multi-source sensing data stream processing is performed to generate sound scene event state characteristics. Combined with urban environmental medium impedance attenuation data, a dynamic virtual cluster topology network is generated, and sound field delay phasor calculation is performed. Control parameters are then distributed to the transducer array for sound suppression.
It enables accurate prediction and effective suppression of sudden acoustic energy leakage events, improves the level of intelligent urban acoustic management, and ensures the long-term stable and tranquil state of the sound environment.
Smart Images

Figure CN122132889A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban soundscape control technology, and in particular to an adaptive control method and system for urban soundscape based on multi-source data. Background Technology
[0002] Urban soundscape adaptive regulation is a multi-dimensional environmental perception and data intervention business under the smart city governance framework. It mainly relies on probes scattered in the blocks to capture the characteristics of underlying noise fluctuations. After aligning with the environmental management tolerance limit database at the core of the network computing server, the control array command parameters are sent down to the front-end control circuit along the street to trigger the loudspeaker transducer to output directional suppression of recoil sound bands. The aim is to use high-precision digital algorithms to allocate and guide physical transducer devices to maintain a long-term stable and quiet state of soundscape environment in a large-scale public grid. When faced with a sudden acoustic energy leakage event with the property of spreading and penetrating in an open square area, the lack of network joint arrangement array data reshaping logic that combines the local topographical obstacle reduction characteristics, and the absence of a joint network control response calculation link that delineates associated allied nodes by energy reduction measure, results in a leak at only one point, which causes the full-dimensional network coverage resistance command to become fragmented and isolated at the landing point and fails to achieve a tight network to contain and suppress the noise. Summary of the Invention
[0003] This invention provides a method and system for adaptive control of urban soundscape based on multi-source data, used for soundscape adjustment processing.
[0004] To achieve the above objectives, the present invention adopts the following technical solution: Firstly, a method for adaptive control of urban soundscape based on multi-source data is provided. This system includes: Acoustic audio data and visual image data synchronously collected by sensing units deployed in urban spatial nodes are acquired and processed by timestamp alignment to generate multi-source sensing data streams; Perform cross-modal alignment calculation and acoustic delay localization spatial matrix measurement at the feature level on the multi-source sensing data stream to generate sound scene event state features that include sound source type classification of sound-emitting entities and absolute three-dimensional coordinate vectors. Obtain a grid node distribution topology map containing physical coordinates and device communication address codes, perform spatial impedance radiation attenuation simulation calculations in conjunction with the sound scene event state characteristics, filter out the node set within the radiation influence boundary, and then generate a dynamic virtual cluster topology network. Based on the classification and positioning attributes in the sound scene event state characteristics and the geometric mapping relationship between each node in the dynamic virtual cluster topology network, the sound field delay phasor calculation is performed, and the sound field collaborative control parameters are distributed to the corresponding node terminals to execute the transducer array sound suppression drive output operation.
[0005] Optionally, the multi-source sensing data stream performs cross-modal alignment calculation and acoustic delay localization spatial matrix measurement at the feature level to generate sound scene event state features including sound source type classification and absolute three-dimensional coordinate vectors, including: The acoustic audio data within the multi-source sensing data stream is subjected to time-frequency domain transformation and feature dimension compression and stripping to generate an audio logarithmic spectrogram vector. Extract the region of interest from the visual image data in the multi-source perception data stream and perform a convolutional sequence texture feature extraction operation to generate a visual spatiotemporally continuous feature vector. A deep neural network weight mapping model is obtained for performing a multidimensional waveform voiceprint classification and identification task. The deep neural network weight mapping model is used to project and map the audio logarithmic spectrogram vector and the visual spatiotemporal continuous feature vector to a hidden layer space with a uniform hidden layer scale, and the feature tensors are added and concatenated to calculate the output of the sound source type classification of the sound-producing entity. The core program of the generalized cross-correlation time difference algorithm is used to process the intercepted acoustic audio data, and combined with reading and extracting the physical offset ranging calibration constant built into the base of the sensing unit for projection calculation to obtain a three-dimensional coordinate vector. Then, the data assembly command is called to bind and associate the sound source type classification of the sound-emitting entity with the absolute three-dimensional coordinate vector.
[0006] Optionally, a grid node distribution topology map containing physical coordinates and device communication addressing codes is obtained. Combined with the soundscape event state characteristics, spatial impedance radiation attenuation simulation calculations are performed to filter out the node set within the radiation influence boundary and generate a dynamic virtual cluster topology network, including: Obtain data from a physical space medium fading distribution layer library containing the assigned impedance attenuation constants of urban environmental buildings and green spaces; Combining the internal environmental constants of the physical space medium fading distribution layer library data, extrapolation and expansion calculations are performed on the absolute three-dimensional coordinate vector within the sound scene event state features, along with spherical sound wave inverse tracking energy reduction and dimensionality reduction simulation derivation calculations, to generate equal sound pressure attenuation envelope edge contour lattice data containing the residual gradient of outward radiated energy. The edge contour lattice data of the equal sound pressure attenuation envelope is projected in the same direction and imported into the grid node distribution topology map for collision superposition matching test operation. By solving the Euclidean distance threshold ranging method, all target hardware trapped in the inner circle of the high-pressure radiation coverage area and the adjacent buffer safety zone of edge decay extension are marked and extracted. The dynamic virtual cluster topology network with exclusive scheduling configuration working communication control cluster ownership relationship only for this target interference task is generated.
[0007] Optionally, based on the classification and positioning attributes in the soundscape event state features and the geometric mapping relationship between each node in the dynamic virtual cluster topology network, sound field delay phasor calculation is performed to generate sound field collaborative control parameters with group role assignment code and underlying phase flip configuration driving code, including: The system extracts the state features of multiple sound scene events from a set of coordinate sliding sequence arrangements within a specific time window slice in a multi-frame arrangement state that is in continuous sampling. The extraction and separation of multiple absolute three-dimensional coordinate vectors, which are encapsulated within the state features of the multiple sound scene events and whose numerical values exhibit continuous wandering and changing characteristics, are performed. Then, second-order partial derivative temporal difference operations are run in the mathematical kernel, combined with the least squares method to solve the trajectory dynamics path penetration geometric fitting equation. This process generates extracted entity instantaneous continuous motion trajectory route prediction flow extension vector line segment vector features that embody the calibration and depiction of the next moment's displacement extension direction and possess predictive guidance movement partial derivative velocity measurement attribute parameters.
[0008] Optionally, sound field delay phasor calculation is performed based on the classification and positioning attributes in the sound scene event state characteristics and the geometric mapping relationship between each node in the dynamic virtual cluster topology network, generating sound field collaborative control parameters with group role assignment code and underlying phase flip configuration driving code.
[0009] Optionally, the monitoring channel continuously captures and collects the status features of the sound scene event, and initiates unpacking detection to scan and detect the internal features of the value exceeding the tolerance warning alarm abnormal boundary touch line status red line marker parameter marker. This content is then assembled, spliced, reprinted, engraved, composite production, and packaged to form a spliced and recombined spatiotemporal segment record block empirical trace content data stream file with playback value and real-time reconstruction function. Through engineering methods that drive the closed loop of evidence data flow from sound scene anomalies, automated evidence collection and assignment for noise issues have been achieved, ensuring cross-domain collaboration from acoustic and physical perception to social governance.
[0010] Optionally, the receiving patrol terminal and feedback terminal use the physical external network channel to submit reports and return records containing grassroots grid supervision and patrol records. The grid security personnel also conduct on-site investigations and record the information by visual inspection and actual hearing tests. It acquires the ability to persist records for a long time without power loss at the system hard disk level, and supports tree-structured network-based query indexes at the underlying level to extract characteristic record history; The program suite of modules for obtaining, installing, mounting, and storing data in a remote backend is used to receive raw waveform feature input, classify, filter, qualitatively judge, and calculate the results. By leveraging the network to complete the brain's evolution and upgrade, a successful intelligent leap in the ability to intervene in the sound system is achieved, enabling the full-stage implementation of intelligent application self-enhancement, including data entry, task execution, and business implementation.
[0011] Optionally, the so-called processing hardware brain control entity, which is embedded in the metal back panel control box of the speaker column rainproof shell, integrates a local area network communication port for communication interaction, and has built-in decoding protocol, is specifically responsible for turning compressed instructions into microcontroller recognition, and has the function of signal processing, operation logic scheduling center. This hardware name symbol characteristics refer to the array control center with mathematical processing core, that is, the digital signal processing control main core control circuit board integrated module component.
[0012] Optionally, the storage and attachment include the carrying capacity, and the use of different attribute positioning base maps and time (sunrise, sunset, get off work hours, tide peaks and troughs) are used to define the area through the city's long-term development plan surveying and mapping. By combining packet download with the autonomous operation of local computing power, the effectiveness of the urban acoustic governance network in suppressing diffuse and transient sound sources has been enhanced.
[0013] Secondly, an adaptive soundscape control system for cities based on multi-source data includes: The perception and edge synchronization module is used to acquire acoustic audio data and visual image data synchronously collected by perception units deployed in urban spatial nodes, and generate multi-source perception data streams through timestamp alignment processing. The event parsing and calibration module is used to perform cross-modal alignment calculation and acoustic delay localization spatial matrix calculation at the feature level on the multi-source sensing data stream, and generate sound scene event state features that include sound source type classification of sound-emitting entities and absolute three-dimensional coordinate vectors. The topology filtering and networking module is used to obtain a grid node distribution topology map containing physical coordinates and device communication address codes, and perform spatial impedance radiation attenuation simulation calculations in combination with the sound scene event state characteristics to filter out the node set within the radiation influence boundary and then generate a dynamic virtual cluster topology network. The delay phasor calculation and distribution module is used to perform sound field delay phasor calculation based on the classification and positioning attributes in the sound scene event state characteristics and the geometric mapping relationship between each node in the dynamic virtual cluster topology network, generate sound field collaborative control parameters with group role assignment code and underlying phase flip configuration driving code, and distribute the sound field collaborative control parameters to the corresponding node terminal to execute the transducer array sound suppression drive output operation.
[0014] Thirdly, an electronic device is provided, comprising: a processor and a memory; the memory is used to store a computer program, which, when executed by the processor, causes the electronic device to perform the urban soundscape adaptive control method and system based on multi-source data as described in the first aspect.
[0015] In one possible design, the electronic device described in the third aspect may further include a transceiver. This transceiver may be a transceiver circuit or an interface circuit. The transceiver can be used for communication between the electronic device described in the third aspect and other electronic devices.
[0016] In the embodiments of the present invention, the electronic device described in the third aspect may be a terminal, or a chip (system) or other component or assembly disposed in the terminal, or a system containing the terminal.
[0017] Fourthly, a computer-readable storage medium is provided, comprising: a computer program or instructions; when the computer program or instructions are executed on a computer, the computer causes the computer to perform the urban soundscape adaptive control method and system based on multi-source data as described in the first aspect.
[0018] In summary, the above methods and systems have the following technical effects: This invention overcomes the shortcomings of single environmental sound sensors, such as misjudgment and coordinate blind spots, caused by complex background noise from human voices and traffic flow, by constructing a hidden layer alignment logic architecture that integrates the underlying audio logarithmic spectrum and visual spatiotemporal continuous features with a multimodal tensor addition layer. Furthermore, it creatively replaces blind network broadcasting with digital mapping simulation addressing by performing spatial spherical wave back-radiation energy subtraction model simulation calculations using urban dielectric impedance fading constant distribution layer data. Subsequently, by extracting features from the flowing spatial three-dimensional absolute coordinate matrix sequence data frames and performing second-order partial derivative and fitting pathfinding operations, it elevates the original passive response process to instantaneously random passing sound-emitting objects to a prospective interferometric inference dimension that can obtain and predict extended displacement spatial pointing vector spectrum in advance. Finally, it provides a network-deep cognitive iterative supplementary digital computation evolution architecture structure that removes complex measurement and constraint systems and offsets leakage-proof multidimensional defense factors by setting proportional adjustment multiples. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of a control system provided in an embodiment of the present invention. Detailed Implementation
[0020] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0021] In this embodiment of the invention, "instruction" can include direct and indirect instructions, as well as explicit and implicit instructions. The information indicated by a certain piece of information is called the information to be instructed. In specific implementation, there are many ways to instruct the information to be instructed, such as, but not limited to, directly instructing the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly instruct the information to be instructed by instructing other information, where there is a correlation between the other information and the information to be instructed. It can also instruct only a part of the information to be instructed, while the other parts are known or pre-agreed upon. For example, the instruction of specific information can be achieved by using a pre-agreed (e.g., protocol-defined) arrangement of various pieces of information, thereby reducing instruction overhead to some extent. Simultaneously, common parts of various pieces of information can be identified and uniformly indicated to reduce the instruction overhead caused by individually indicating the same information.
[0022] Furthermore, the specific indication method can also be any existing indication method, such as, but not limited to, the above-mentioned indication methods and their various combinations. Specific details of various indication methods can be found in existing technologies, and will not be elaborated upon here. As described above, for example, when multiple pieces of information of the same type need to be indicated, the indication methods for different pieces of information may differ. In specific implementation, the required indication method can be selected according to specific needs. This embodiment of the invention does not limit the selected indication method; therefore, the indication methods involved in this embodiment of the invention should be understood to cover various methods that enable the party to be indicated to obtain the information to be indicated.
[0023] It should be understood that the information to be indicated can be sent as a whole or divided into multiple sub-information messages sent separately, and the sending period and / or timing of these sub-information messages can be the same or different. The specific sending method is not limited in this embodiment of the invention. The sending period and / or timing of these sub-information messages can be predefined, for example, according to a protocol, or configured by the sending device by sending configuration information to the receiving device.
[0024] "Predefined" or "pre-configured" can be achieved by pre-saving corresponding codes, tables, or other means that can be used to indicate relevant information in the device. This embodiment of the invention does not limit the specific implementation method. "Saving" can refer to saving in one or more memories. These memories can be separate installations or integrated into the encoder, decoder, processor, or electronic device. Alternatively, some memories can be separately installed, while others are integrated into the decoder, processor, or electronic device. The type of memory can be any form of storage medium, and this embodiment of the invention does not limit this.
[0025] In the embodiments of this invention, the “protocol” may refer to a protocol family in the field of communication, a standard protocol with a similar protocol family frame structure, or a related protocol applied to a future urban soundscape adaptive control method and system based on multi-source data. The embodiments of this invention do not specifically limit this.
[0026] In this embodiment of the invention, descriptions such as "when," "under the circumstances," "if," and "if" all refer to the device making corresponding processing under certain objective circumstances, and are not limited to a specific time. They do not require the device to make a judgment action during implementation, nor do they imply any other limitations.
[0027] In the description of the embodiments of the present invention, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can represent A or B. "And / or" in the embodiments of the present invention is merely a description of the relationship between the related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Furthermore, in the description of the embodiments of the present invention, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. Additionally, to facilitate a clear description of the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or order of execution, and that "first," "second," etc., are not necessarily different. Furthermore, in the embodiments of this invention, words such as "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this invention should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner for ease of understanding.
[0028] The network architecture and business scenarios described in the embodiments of this invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems.
[0029] The above combination Figure 1 The method provided by the embodiments of the present invention is described in detail below. A method for adaptive control of urban soundscape based on multi-source data, used to perform the method provided by the embodiments of the present invention, is described in detail below. The apparatus includes: Acoustic audio data and visual image data synchronously collected by sensing units deployed in urban spatial nodes are acquired and processed by timestamp alignment to generate multi-source sensing data streams; Perform cross-modal alignment calculation and acoustic delay localization spatial matrix measurement at the feature level on the multi-source sensing data stream to generate sound scene event state features that include sound source type classification of sound-emitting entities and absolute three-dimensional coordinate vectors. Obtain a grid node distribution topology map containing physical coordinates and device communication address codes, perform spatial impedance radiation attenuation simulation calculations in conjunction with the sound scene event state characteristics, filter out the node set within the radiation influence boundary, and then generate a dynamic virtual cluster topology network. Based on the classification and positioning attributes in the sound scene event state characteristics and the geometric mapping relationship between each node in the dynamic virtual cluster topology network, the sound field delay phasor calculation is performed, and the sound field collaborative control parameters are distributed to the corresponding node terminals to execute the transducer array sound suppression drive output operation.
[0030] Optionally, the multi-source sensing data stream performs cross-modal alignment calculation and acoustic delay localization spatial matrix measurement at the feature level to generate sound scene event state features including sound source type classification and absolute three-dimensional coordinate vectors, including: The acoustic audio data within the multi-source sensing data stream is subjected to time-frequency domain transformation and feature dimension compression and stripping to generate an audio logarithmic spectrogram vector. Extract the region of interest from the visual image data in the multi-source perception data stream and perform a convolutional sequence texture feature extraction operation to generate a visual spatiotemporally continuous feature vector. A deep neural network weight mapping model is obtained for performing a multidimensional waveform voiceprint classification and identification task. The deep neural network weight mapping model is used to project and map the audio logarithmic spectrogram vector and the visual spatiotemporal continuous feature vector to a hidden layer space with a uniform hidden layer scale, and the feature tensors are added and concatenated to calculate the output of the sound source type classification of the sound-producing entity. The core program of the generalized cross-correlation time difference algorithm is used to process the intercepted acoustic audio data, and combined with reading and extracting the physical offset ranging calibration constant built into the base of the sensing unit for projection calculation to obtain a three-dimensional coordinate vector. Then, the data assembly command is called to bind and associate the sound source type classification of the sound-emitting entity with the absolute three-dimensional coordinate vector.
[0031] Optionally, a grid node distribution topology map containing physical coordinates and device communication addressing codes is obtained. Combined with the soundscape event state characteristics, spatial impedance radiation attenuation simulation calculations are performed to filter out the node set within the radiation influence boundary and generate a dynamic virtual cluster topology network, including: Obtain data from a physical space medium fading distribution layer library containing the assigned impedance attenuation constants of urban environmental buildings and green spaces; Combining the internal environmental constants of the physical space medium fading distribution layer library data, extrapolation and expansion calculations are performed on the absolute three-dimensional coordinate vector within the sound scene event state features, along with spherical sound wave inverse tracking energy reduction and dimensionality reduction simulation derivation calculations, to generate equal sound pressure attenuation envelope edge contour lattice data containing the residual gradient of outward radiated energy. The edge contour lattice data of the equal sound pressure attenuation envelope is projected in the same direction and imported into the grid node distribution topology map for collision superposition matching test operation. By solving the Euclidean distance threshold ranging method, all target hardware trapped in the inner circle of the high-pressure radiation coverage area and the adjacent buffer safety zone of edge decay extension are marked and extracted. The dynamic virtual cluster topology network with exclusive scheduling configuration working communication control cluster ownership relationship only for this target interference task is generated.
[0032] Optionally, based on the classification and positioning attributes in the soundscape event state features and the geometric mapping relationship between each node in the dynamic virtual cluster topology network, sound field delay phasor calculation is performed to generate sound field collaborative control parameters with group role assignment code and underlying phase flip configuration driving code, including: The system extracts the state features of multiple sound scene events from a set of coordinate sliding sequence arrangements within a specific time window slice in a multi-frame arrangement state that is in continuous sampling. The extraction and separation of multiple absolute three-dimensional coordinate vectors, which are encapsulated within the state features of the multiple sound scene events and whose numerical values exhibit continuous wandering and changing characteristics, are performed. Then, second-order partial derivative temporal difference operations are run in the mathematical kernel, combined with the least squares method to solve the trajectory dynamics path penetration geometric fitting equation. This process generates extracted entity instantaneous continuous motion trajectory route prediction flow extension vector line segment vector features that embody the calibration and depiction of the next moment's displacement extension direction and possess predictive guidance movement partial derivative velocity measurement attribute parameters.
[0033] Optionally, sound field delay phasor calculation is performed based on the classification and positioning attributes in the sound scene event state characteristics and the geometric mapping relationship between each node in the dynamic virtual cluster topology network, generating sound field collaborative control parameters with group role assignment code and underlying phase flip configuration driving code.
[0034] Optionally, the monitoring channel continuously captures and collects the status features of the sound scene event, and initiates unpacking detection to scan and detect the internal features of the value exceeding the tolerance warning alarm abnormal boundary touch line status red line marker parameter marker. This content is then assembled, spliced, reprinted, engraved, composite production, and packaged to form a spliced and recombined spatiotemporal segment record block empirical trace content data stream file with playback value and real-time reconstruction function. Through engineering methods that drive the closed loop of evidence data flow from sound scene anomalies, automated evidence collection and assignment for noise issues have been achieved, ensuring cross-domain collaboration from acoustic and physical perception to social governance.
[0035] Optionally, the receiving patrol terminal and feedback terminal use the physical external network channel to submit reports and return records containing grassroots grid supervision and patrol records. The grid security personnel also conduct on-site investigations and record the information by visual inspection and actual hearing tests. It acquires the ability to persist records for a long time without power loss at the system hard disk level, and supports tree-structured network-based query indexes at the underlying level to extract characteristic record history; The program suite of modules for obtaining, installing, mounting, and storing data in a remote backend is used to receive raw waveform feature input, classify, filter, qualitatively judge, and calculate the results. By leveraging the network to complete the brain's evolution and upgrade, a successful intelligent leap in the ability to intervene in the sound system is achieved, enabling the full-stage implementation of intelligent application self-enhancement, including data entry, task execution, and business implementation.
[0036] Optionally, the so-called processing hardware brain control entity, which is embedded in the metal back panel control box of the speaker column rainproof shell, integrates a local area network communication port for communication interaction, and has built-in decoding protocol, is specifically responsible for turning compressed instructions into microcontroller recognition, and has the function of signal processing, operation logic scheduling center. This hardware name symbol characteristics refer to the array control center with mathematical processing core, that is, the digital signal processing control main core control circuit board integrated module component.
[0037] Optionally, the storage and attachment include the carrying capacity, and the use of different attribute positioning base maps and time (sunrise, sunset, get off work hours, tide peaks and troughs) are used to define the area through the city's long-term development plan surveying and mapping. By combining packet download with the autonomous operation of local computing power, the effectiveness of the urban acoustic governance network in suppressing diffuse and transient sound sources has been enhanced.
[0038] This invention provides a schematic diagram of the structure of an electronic device. Exemplarily, the electronic device can be a network device, or a chip (system) or other component or assembly that can be disposed in a network device. As shown, the electronic device may include a processor. Optionally, the electronic device may also include a memory and / or a transceiver. The processor is coupled to the memory and transceiver, for example, by means of a communication bus connection.
[0039] The following section provides a detailed introduction to each component of the electronic device, with reference to the diagram: In this context, the processor is the control center of the electronic device. It can be a single processor or a collective term for multiple processing elements. For example, a processor can be one or more central processing units (CPUs), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0040] Optionally, the processor can perform various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory, such as executing the urban soundscape adaptive control method and system based on multi-source data shown in the figure above.
[0041] In a specific implementation, as one example, the processor may include one or more CPUs, such as the CPU and CPU shown in the figure.
[0042] In a specific implementation, as one example, the electronic device may also include multiple processors. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, a processor may refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).
[0043] The memory is used to store the software program that executes the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.
[0044] Optionally, the memory can be read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory can be integrated with the processor or exist independently and coupled to the processor through the interface circuit of the electronic device; the embodiments of the present invention do not specifically limit this.
[0045] A transceiver is used for communication with other electronic devices. For example, if the electronic device is a terminal, the transceiver can be used to communicate with a network device or with another terminal device. As another example, if the electronic device is a network device, the transceiver can be used to communicate with a terminal or with another network device.
[0046] Optionally, the transceiver may include a receiver and a transmitter (not shown separately in the figure). The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0047] Optionally, the transceiver can be integrated with the processor or exist independently and coupled to the processor through the interface circuit of the electronic device. This embodiment of the invention does not specifically limit this.
[0048] It is understood that the structure of the electronic device shown in the figure does not constitute a limitation on the electronic device. The actual electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0049] Furthermore, the technical effects of the electronic devices can be referred to in the above-described method embodiments for the adaptive control method and system for urban soundscape based on multi-source data, which will not be elaborated here.
[0050] It should be understood that the processor in the embodiments of the present invention can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0051] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0052] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0053] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0054] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0055] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0056] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0057] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0058] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0059] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0060] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0061] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0062] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for adaptive control of urban soundscape based on multi-source data, characterized in that, Includes the following steps: Acoustic audio data and visual image data synchronously collected by sensing units deployed in urban spatial nodes are acquired and processed by timestamp alignment to generate multi-source sensing data streams; Perform cross-modal alignment calculation and acoustic delay localization spatial matrix measurement at the feature level on the multi-source sensing data stream to generate sound scene event state features that include sound source type classification of sound-emitting entities and absolute three-dimensional coordinate vectors. Obtain a grid node distribution topology map containing physical coordinates and device communication address codes, perform spatial impedance radiation attenuation simulation calculations in conjunction with the sound scene event state characteristics, filter out the node set within the radiation influence boundary, and then generate a dynamic virtual cluster topology network. Based on the classification and positioning attributes in the sound scene event state characteristics and the geometric mapping relationship between each node in the dynamic virtual cluster topology network, the sound field delay phasor calculation is performed, and the sound field collaborative control parameters are distributed to the corresponding node terminals to execute the transducer array sound suppression drive output operation.
2. The urban soundscape adaptive control method based on multi-source data as described in claim 1, characterized in that, The multi-source sensing data stream performs cross-modal alignment calculations and acoustic delay localization spatial matrix measurements at the feature level, generating sound scene event state features that include sound source type classification and absolute three-dimensional coordinate vectors, including: The acoustic audio data within the multi-source sensing data stream is subjected to time-frequency domain transformation and feature dimension compression and stripping to generate an audio logarithmic spectrogram vector. Extract the region of interest from the visual image data in the multi-source perception data stream and perform a convolutional sequence texture feature extraction operation to generate a visual spatiotemporally continuous feature vector. A deep neural network weight mapping model is obtained for performing a multidimensional waveform voiceprint classification and identification task. The deep neural network weight mapping model is used to project and map the audio logarithmic spectrogram vector and the visual spatiotemporal continuous feature vector to a hidden layer space with a uniform hidden layer scale, and the feature tensors are added and concatenated to calculate the output of the sound source type classification of the sound-producing entity. The core program of the generalized cross-correlation time difference algorithm is used to process the intercepted acoustic audio data, and combined with reading and extracting the physical offset ranging calibration constant built into the base of the sensing unit for projection calculation to obtain a three-dimensional coordinate vector. Then, the data assembly command is called to bind and associate the sound source type classification of the sound-emitting entity with the absolute three-dimensional coordinate vector.
3. The urban soundscape adaptive control method based on multi-source data as described in claim 1, characterized in that, Obtain a grid node distribution topology map containing physical coordinates and device communication addressing codes. Combine this with the spatial impedance radiation attenuation simulation calculation based on the sound scene event state characteristics. Select the node set within the radiation influence boundary to generate a dynamic virtual cluster topology network, including: Obtain data from a physical space medium fading distribution layer library containing the assigned impedance attenuation constants of urban environmental buildings and green spaces; Combining the internal environmental constants of the physical space medium fading distribution layer library data, extrapolation and expansion calculations are performed on the absolute three-dimensional coordinate vector within the sound scene event state features, along with spherical sound wave inverse tracking energy reduction and dimensionality reduction simulation derivation calculations, to generate equal sound pressure attenuation envelope edge contour lattice data containing the residual gradient of outward radiated energy. The edge contour lattice data of the equal sound pressure attenuation envelope is projected in the same direction and imported into the grid node distribution topology map for collision superposition matching test operation. By solving the Euclidean distance threshold ranging method, all target hardware trapped in the inner circle of the high-pressure radiation coverage area and the adjacent buffer safety zone of edge decay extension are marked and extracted. The dynamic virtual cluster topology network with exclusive scheduling configuration working communication control cluster ownership relationship only for this target interference task is generated.
4. The urban soundscape adaptive control method based on multi-source data as described in claim 1, characterized in that, Based on the classification and positioning attributes in the soundscape event state features and the geometric mapping relationship between each node in the dynamic virtual cluster topology network, sound field delay phasor calculation is performed to generate sound field collaborative control parameters with group role assignment code and underlying phase flip configuration driving code, including: The system extracts the state features of multiple sound scene events from a set of coordinate sliding sequence arrangements within a specific time window slice in a multi-frame arrangement state that is in continuous sampling. The extraction and separation of multiple absolute three-dimensional coordinate vectors, which are encapsulated within the state features of the multiple sound scene events and whose numerical values exhibit continuous wandering and changing characteristics, are performed. Then, second-order partial derivative temporal difference operations are run in the mathematical kernel, combined with the least squares method to solve the trajectory dynamics path penetration geometric fitting equation. This process generates extracted entity instantaneous continuous motion trajectory route prediction flow extension vector line segment vector features that embody the calibration and depiction of the next moment's displacement extension direction and possess predictive guidance movement partial derivative velocity measurement attribute parameters.
5. The urban soundscape adaptive control method based on multi-source data as described in claim 1, characterized in that, Based on the classification and positioning attributes in the sound scene event state characteristics and the geometric mapping relationship between each node in the dynamic virtual cluster topology network, sound field delay phasor calculation is performed to generate sound field collaborative control parameters with group role assignment code and underlying phase flip configuration driving code.
6. The urban soundscape adaptive control method based on multi-source data as described in claim 1, characterized in that, The monitoring channel continuously captures and receives the status features of the sound scene event, and initiates unpacking and detection. It scans and detects the internal features of the value exceeding the tolerance, warning, alarm, abnormal boundary, touch line, status red line, marker, parameter marker, and other features. This content is then assembled, spliced, reprinted, engraved, composite, produced, and packaged to form a spliced and recombined spatiotemporal segment record block empirical trace content data stream file with playback value and real-time reconstruction function. Through engineering methods that drive the closed loop of evidence data flow from sound scene anomalies, automated evidence collection and assignment for noise issues have been achieved, ensuring cross-domain collaboration from acoustic and physical perception to social governance.
7. The urban soundscape adaptive control method based on multi-source data as described in claim 1, characterized in that, The system receives patrol terminals and feedback terminals, uses physical external network channels to submit reports and return records containing grassroots grid supervision and patrol records, and allows grid security personnel to conduct on-site investigations and record the findings through physical visual inspection and actual auditory testing. It acquires the ability to persist records for a long time without power loss at the system hard disk level, and supports tree-structured network-based query indexes at the underlying level to extract characteristic record history; The program suite of modules for obtaining, installing, mounting, and storing data in a remote backend is used to receive raw waveform feature input, classify, filter, qualitatively judge, and calculate the results. By leveraging the network to complete the brain's evolution and upgrade, a successful intelligent leap in the ability to intervene in the sound system is achieved, enabling the full-stage implementation of intelligent application self-enhancement, including data entry, task execution, and business implementation.
8. The urban soundscape adaptive control method based on multi-source data as described in claim 1, characterized in that, The so-called processing hardware brain control entity, which is embedded in the metal back panel of the speaker column's rainproof shell, has a local area network communication port for communication and interaction. It also has a built-in decoding protocol and is specifically responsible for converting compressed instructions into something that a microcontroller can understand. This hardware entity has the function of a signal processing, operation logic, and scheduling center. The hardware name symbol characteristics refer to the array control center with a mathematical processing core, namely the digital signal processing control main core control circuit board integrated module component.
9. The urban soundscape adaptive control method based on multi-source data as described in claim 1, characterized in that, The acquisition of storage and attachment includes the carrying capacity, and the use of different attribute positioning base maps and time (sunrise, sunset, get off work hours, tide peaks and troughs) to define the area through the city's long-term development plan and map surveying. By combining packet download with the autonomous operation of local computing power, the effectiveness of the urban acoustic governance network in suppressing diffuse and transient sound sources has been enhanced.
10. A multi-source data-based adaptive urban soundscape control system, applied to the multi-source data-based adaptive urban soundscape control method according to any one of claims 1-9, characterized in that: The perception and edge synchronization module is used to acquire acoustic audio data and visual image data synchronously collected by perception units deployed in urban spatial nodes, and generate multi-source perception data streams through timestamp alignment processing. The event parsing and calibration module is used to perform cross-modal alignment calculation and acoustic delay localization spatial matrix calculation at the feature level on the multi-source sensing data stream, and generate sound scene event state features that include sound source type classification of sound-emitting entities and absolute three-dimensional coordinate vectors. The topology filtering and networking module is used to obtain a grid node distribution topology map containing physical coordinates and device communication address codes, and perform spatial impedance radiation attenuation simulation calculations in combination with the sound scene event state characteristics to filter out the node set within the radiation influence boundary and then generate a dynamic virtual cluster topology network. The delay phasor calculation and distribution module is used to perform sound field delay phasor calculation based on the classification and positioning attributes in the sound scene event state characteristics and the geometric mapping relationship between each node in the dynamic virtual cluster topology network, generate sound field collaborative control parameters with group role assignment code and underlying phase flip configuration driving code, and distribute the sound field collaborative control parameters to the corresponding node terminal to execute the transducer array sound suppression drive output operation.