Visual communication advertisement design system and method
By collecting light and voiceprint data through the IoT sensor network to generate an environmental perception state tensor, combined with a dynamic lighting matrix and acoustic beam width angle, the advertisement can be rendered synchronously with sound and light. This solves the problems of unclear images and inaccurate audio transmission in different lighting scenarios, and improves the rendering effect of advertisements.
Patent Information
- Application Number
- CN202510793799.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing advertising rendering technology is unable to dynamically adjust lighting effects and audio transmission according to ambient light and sound, resulting in unclear or dazzling advertisements in different lighting scenarios, and the inability of audio to be accurately and targetedly transmitted, unable to effectively reach the target audience.
Light intensity and environmental soundprint data are collected through the IoT sensor network to generate an environmental perception state tensor. Dynamic rendering is performed by combining light and audio parameters to generate a dynamic lighting matrix and acoustic beam width angle to achieve synchronous sound and light rendering.
It enhances the temporal and spatial consistency between advertising content and the real environment, improves the realism and environmental adaptability of picture rendering, and ensures that advertisements are naturally integrated into complex scenes and effectively convey information.
Smart Images

Figure CN120655355A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of advertising rendering technology, and more specifically, to a visual communication advertising design system and method. Background Art
[0002] In today's digital age, most advertisements are delivered based on user portraits and the characteristics of the delivery platform. With the rapid development of Internet technology, advertising delivery scenarios have become increasingly diversified, and people's requirements for the presentation effects of advertisements are increasing. They expect advertisements to be deeply integrated with the environment and bring a more natural and comfortable visual and auditory experience.
[0003] Existing advertising rendering technology makes it difficult to dynamically adjust the lighting effects of advertising images according to ambient lighting factors in visual presentation, resulting in problems such as advertisements being too dark to see or too bright and dazzling in different lighting scenarios. In terms of auditory presentation, audio playback is mostly fixed in mode, and cannot achieve precise and targeted transmission based on ambient sounds and audience positions. It is easily masked by ambient noise, resulting in the advertising content being unable to accurately reach the target audience, making the attractiveness and influence of the advertising content unable to achieve the expected results. Therefore, how to perform sound and light coordinated rendering of target advertisements and dynamically adapt sound and light parameters to achieve targeted transmission to the audience has become a difficult problem faced by the industry. Summary of the Invention
[0004] The present application provides a visual communication advertising design system and method, which can perform sound and light collaborative rendering and dynamic adaptation of sound and light parameters on target advertisements to achieve targeted dissemination to the audience.
[0005] In a first aspect, the present application provides an advertisement rendering method based on sound and light synergy and environmental perception, for dynamically rendering a designed advertisement in a visual communication advertisement design system, the method comprising: Synchronously collect light intensity data and environmental voiceprint data of advertising scenes through the IoT sensor network; Extracting background voiceprint features of the advertising delivery scene from the environmental voiceprint data, and then performing cross-modal feature alignment on the background voiceprint features and the frequency domain features of light intensity in the light intensity data to obtain an environmental perception state tensor of the advertising delivery scene; Generate a dynamic lighting matrix for rendering the target advertisement by combining the spectral irradiance distribution in the environmental perception state tensor with the anisotropic highlight coefficients of each sub-region in the target advertisement screen; Acquiring spatial depth data of an advertisement delivery scene, identifying the audience's position based on the spatial depth data and the audience's movement trajectory in the target advertisement delivery scene, obtaining the audience spatial domain where the target advertisement is played, and then performing spatial audio redirection on the target advertisement based on the audience spatial domain and the sound field energy distribution in the environmental perception state tensor to obtain the acoustic beamwidth angle when rendering the target advertisement; The advertisement content of the target advertisement is dynamically rendered in a sound and light synchronous manner according to the dynamic illumination matrix and the acoustic beam width angle.
[0006] In some embodiments, extracting background voiceprint features of the advertising delivery scene from the environmental voiceprint data specifically includes: Performing a short-time Fourier transform on the environmental voiceprint data to obtain a sound spectrum of the environmental voiceprint data; Mapping the environmental voiceprint data into a sound field static texture based on the sound spectrogram to obtain a voiceprint sub-feature of each signal vector in the environmental voiceprint data; All voiceprint sub-features are reconstructed by dimensionality reduction to obtain the background voiceprint features of the advertising scene.
[0007] In some embodiments, performing cross-modal feature alignment on the background voiceprint feature and the frequency domain feature of the light intensity in the light intensity data to obtain the environment perception state tensor of the advertising delivery scene specifically includes: extracting frequency domain features of light intensity from the light intensity data; Performing time domain registration on the background voiceprint features and the frequency domain features to obtain the voiceprint-illumination dual-mode features of the advertising delivery scene; The voiceprint-illumination dual-mode features are mapped to a three-dimensional state space using a tensor fusion mechanism to form an environmental perception state tensor of the advertising delivery scenario.
[0008] In some embodiments, generating a dynamic lighting matrix for rendering a target advertisement by combining the spectral irradiance distribution in the environmental perception state tensor with the anisotropic highlight coefficients of each sub-region in the target advertisement screen specifically includes: Divide the target advertising screen into sub-regions to obtain multiple sub-regions; Extracting the anisotropic highlight coefficient of each sub-region in the target advertising image; determining an environment matching factor for each sub-region based on the spectral irradiance distribution in the environment perception state tensor; The environment matching factors and anisotropic highlight coefficients of all sub-areas are mapped to the lighting rendering parameter space to generate a dynamic lighting matrix for rendering the target advertisement.
[0009] In some embodiments, identifying the audience position based on the spatial depth data and the audience movement trajectory in the target advertisement delivery scene to obtain the audience spatial domain for the target advertisement playback specifically includes: generating a three-dimensional point cloud map of the target advertising delivery scene based on the spatial depth data; Collect millimeter-wave radar signals of target advertising scenes and extract audience movement trajectories; Performing spatial fusion registration on the three-dimensional point cloud image and the audience's movement trajectory to obtain the spatial position distribution of the target audience; Sub-areas are divided based on the spatial position distribution to obtain the audience spatial domain for target advertisement broadcasting.
[0010] In some embodiments, performing spatial audio redirection on a target advertisement based on the audience spatial domain and the sound field energy distribution in the environmental perception state tensor to obtain an acoustic beamwidth angle when rendering the target advertisement specifically includes: Determining an auditory pointing vector for each audience member based on audience distribution and direction information in the audience space domain; Generate the audience area coverage angle of the target advertising scene through all auditory pointing vectors; Adjusting the frequency domain of the main audio channel according to the sound field energy distribution in the environmental perception state tensor to obtain an audio projection angle; An acoustic beamwidth angle for rendering a target advertisement is generated based on the audience area coverage angle and the audio projection angle.
[0011] In some embodiments, performing acoustic and optical synchronized dynamic rendering of the advertisement content of the target advertisement according to the dynamic illumination matrix and the acoustic beamwidth angle specifically includes: Adjusting camera parameters in a rendering engine pipeline based on the dynamic lighting matrix to achieve dynamic lighting rendering; adjusting the amplitude and phase of the multi-channel audio according to the acoustic beamwidth angle; A time synchronization lock mechanism is constructed to control the synchronization error between the visual rendering and audio rendering of the target advertising content, and then the advertising content of the target advertisement is dynamically rendered in sound and light synchronization according to the synchronization error.
[0012] In a second aspect, the present application provides a visual communication advertising design system, the system including an advertising rendering unit, the advertising rendering unit including: The acquisition module is used to synchronously collect light intensity data and environmental voiceprint data of the advertising scene through the IoT sensor network; a processing module configured to extract background voiceprint features of the advertising delivery scene from the environmental voiceprint data, and then perform cross-modal feature alignment on the background voiceprint features and the frequency domain features of light intensity in the light intensity data to obtain an environmental perception state tensor of the advertising delivery scene; The processing module is configured to generate a dynamic lighting matrix for rendering the target advertisement by combining the spectral irradiance distribution in the environmental perception state tensor with the anisotropic highlight coefficients of each sub-region in the target advertisement screen; The processing module is configured to obtain spatial depth data of an advertisement delivery scene, identify an audience position based on the spatial depth data and an audience movement trajectory of the target advertisement delivery scene, obtain an audience spatial domain for the target advertisement, and then perform spatial audio redirection on the target advertisement based on the audience spatial domain and the sound field energy distribution in the environmental perception state tensor to obtain an acoustic beamwidth angle when rendering the target advertisement; An execution module is configured to perform sound and light synchronous dynamic rendering on the advertisement content of the target advertisement according to the dynamic illumination matrix and the acoustic beam width angle.
[0013] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores code, and the processor is configured to obtain the code and execute the above-mentioned advertising rendering method based on sound-light collaboration and environmental perception.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned advertising rendering method based on sound-light collaboration and environmental perception.
[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: In the visual communication advertising design system and method provided in the present application, the light intensity data and environmental soundprint data of the advertising delivery scene are first synchronously collected through the Internet of Things sensor network; the background soundprint features of the advertising delivery scene are extracted from the environmental soundprint data, and then the background soundprint features and the frequency domain features of the light intensity in the light intensity data are cross-modally aligned to obtain the environmental perception state tensor of the advertising delivery scene; the spectral irradiance distribution in the environmental perception state tensor is combined with the anisotropic highlight coefficient of each sub-area in the target advertising picture to generate a dynamic lighting matrix for rendering the target advertisement; the spatial depth data of the advertising delivery scene is obtained, the audience position is identified according to the spatial depth data and the audience movement trajectory of the target advertising delivery scene, and the audience spatial domain of the target advertisement is obtained, and then the spatial audio of the target advertisement is spatially redirected based on the audience spatial domain and the sound field energy distribution in the environmental perception state tensor to obtain the acoustic beam width angle when rendering the target advertisement; the advertising content of the target advertisement is dynamically rendered with sound and light synchronization according to the dynamic lighting matrix and the acoustic beam width angle.
[0016] It can be seen that the present application performs sound and light synchronous dynamic rendering of the advertising content of the target advertisement according to the dynamic lighting matrix and the acoustic beam width angle; first, determining the environmental perception state tensor can obtain a high-dimensional data structure that characterizes the spatiotemporal correlation characteristics and dynamic change laws of sound and light in the advertising delivery scene. The environmental perception state tensor converts the sound and light information in the environment into machine-understandable structured data through cross-modal feature alignment and tensor fusion, which can provide multi-dimensional environmental state input for advertising rendering, so that the system can realize fine adjustment of sound and light parameters according to the coordinated changes of scene sound and light, solve the "one-size-fits-all" rendering defects in traditional technology for dynamic rendering, enhance the spatiotemporal consistency between advertising content and the real environment, and improve the realism and environmental adaptability of picture rendering; then, determining the dynamic lighting matrix can obtain a real-time lighting transformation matrix that integrates the dynamic characteristics of environmental lighting and the reflection parameters of the advertising picture material. The illumination matrix constructs a dynamic mapping relationship between the advertising image and the real scene by integrating the ambient lighting and the material characteristics of the advertising image. The advertising image can be accurately mapped to the three-dimensional scene for lighting consistency rendering, which solves the core pain point of the mismatch between the brightness of the image and the environment in the existing technology, and achieves a natural rendering effect with no exposure in bright areas and details in dark areas. Finally, the acoustic beam width angle is determined to obtain the horizontal angle range that the directional audio emitted from the sound source should cover. The acoustic beam width angle balances the coverage range through adjustable angle parameters, which can balance the audience coverage range and the accuracy of audio transmission, providing a quantitative basis for advertising audio rendering, so that the advertising sound can adapt to the dynamic scene, which can not only cover the target audience group, but also reduce the sound diffusion loss, and improve the overall effect of the advertising sound and light synchronous rendering. In summary, based on the above scheme, the target advertisement can be rendered with sound and light in a coordinated manner and the sound and light parameters can be dynamically adapted to achieve directional transmission to the audience. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is an exemplary flow chart of an advertisement rendering method based on sound and light collaboration and environmental perception according to some embodiments of the present application; Figure 2 is an operational flow chart of determining an environment perception state tensor according to some embodiments of the present application; Figure 3 is an exemplary flow chart of determining a dynamic lighting matrix according to some embodiments of the present application; Figure 4 is a schematic structural diagram of an advertisement rendering unit according to some embodiments of the present application; Figure 5 This is a diagram of the internal structure of a computer device that implements an advertisement rendering method based on sound and light coordination and environmental perception according to some embodiments of the present application. DETAILED DESCRIPTION
[0018] In order to better understand the technical solution of the present application, the technical solution of the present application will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0019] refer to Figure 1 , which is an exemplary flow chart of an advertisement rendering method based on sound-light collaboration and environmental perception according to some embodiments of the present application. The advertisement rendering method based on sound-light collaboration and environmental perception mainly includes the following steps: In step 101, the light intensity data and environmental voiceprint data of the advertising scene are synchronously collected through the Internet of Things sensor network.
[0020] It should be noted that, in this application, light intensity data is time-series perception information that reflects the changes in ambient brightness at different time nodes in the advertising delivery scene. The light intensity data can be used to dynamically evaluate the lighting conditions of the picture to guide visual rendering adaptation; environmental soundprint data refers to the soundprint information with time-frequency characteristics in the background sound signal of the advertising delivery scene. The environmental soundprint data can be used to assist in constructing the spatial sound field and improve the acoustic matching of the advertising content; in specific implementation, the synchronous collection of light intensity data and environmental soundprint data of the advertising delivery scene through the Internet of Things sensor network can be achieved in the following way, namely: a multi-node Internet of Things sensor network can be deployed in the advertising delivery scene, and each Internet of Things sensor node is integrated with a high-sensitivity light sensor (such as a photoresistor or a photodiode module). Block) and directional microphone array, the high-sensitivity light sensor in the IoT sensor node collects the light intensity of the advertising scene at preset collection time intervals (such as 1 second) within a preset collection time period (such as half an hour), and the sequence of all collected light intensities arranged in chronological order is used as the light intensity data of the advertising scene. The directional microphone array in the IoT sensor node then collects the audio signal of the advertising scene within a preset collection time period (such as half an hour), and the collected voiceprint signal is used as the environmental voiceprint data of the advertising scene. In order to ensure the time consistency of data between different nodes, the system adopts the precise time protocol for high-precision time synchronization, so that the collected data has a unified timestamp.
[0021] In step 102, background voiceprint features of the advertising delivery scene are extracted from the environmental voiceprint data, and then cross-modal feature alignment is performed on the background voiceprint features and the frequency domain features of the light intensity in the light intensity data to obtain an environmental perception state tensor of the advertising delivery scene.
[0022] In some embodiments, extracting background voiceprint features of the advertising delivery scene from the environmental voiceprint data can be achieved by using the following steps: Performing a short-time Fourier transform on the environmental voiceprint data to obtain a sound spectrum of the environmental voiceprint data; Mapping the environmental voiceprint data into a sound field static texture based on the sound spectrogram to obtain a voiceprint sub-feature of each signal vector in the environmental voiceprint data; All voiceprint sub-features are reconstructed by dimensionality reduction to obtain the background voiceprint features of the advertising scene.
[0023] In specific implementation, the environmental soundprint data is subjected to a short-time Fourier transform, and the sound spectrum diagram of the environmental soundprint data is obtained in the following manner, namely: first, the environmental sound signal can be converted into an electrical signal by using a microphone, and then the electrical signal is sliced into adjacent signal vectors with time domain overlap through a sliding time window, and each signal vector is subjected to a discrete Fourier transform to obtain the amplitude result corresponding to each signal vector; then, the amplitude result of each signal vector is stacked in the row direction in chronological order to obtain the sound spectrum diagram of the environmental soundprint data; wherein, the sound spectrum diagram is a visual map showing the distribution state of the environmental soundprint data in the time-frequency domain, and the sound spectrum diagram can be used to intuitively reflect the sound intensity at each frequency in different time periods, and provide accurate time-frequency information for subsequent voiceprint sub-feature extraction.
[0024] In specific implementation, the environmental voiceprint data is mapped into a static texture of the sound field based on the sound spectrum graph, and the voiceprint sub-features of each signal vector in the environmental voiceprint data are obtained in the following manner, namely: based on the obtained sound spectrum graph, the spectral energy of each signal vector in the environmental voiceprint data is calculated as the normalized energy coefficient in all frequency bands of the sound spectrum graph (such as using a Mel filter group to compress the 257-dimensional spectral energy into 40-dimensional Mel energy), and the sequence composed of all normalized energy coefficients is input into an existing texture feature extraction tool (such as the texture analysis function of the mathematical software Matlab) to extract the gray-level co-occurrence matrix texture features to obtain the voiceprint sub-features of each signal vector; wherein the voiceprint sub-feature is a time-frequency texture vector representing the local sound field texture, and the voiceprint sub-feature can be used to capture the stable frequency band distribution pattern in the advertising delivery scenario, thereby effectively distinguishing background noise from voice interference.
[0025] It should be noted that, in the present application, the background voiceprint feature is a low-dimensional principal component representation that characterizes the sound stability characteristics of the advertising delivery environment. The background voiceprint feature can be used for subsequent cross-modal alignment and environment modeling to improve the accuracy of cross-domain perception and scene modeling. In specific implementation, all voiceprint sub-features are reconstructed by dimensionality reduction to obtain the background voiceprint feature of the advertising delivery scene. This can be achieved in the following way: the voiceprint sub-features of all signal vectors can be input into a pre-trained encoder model (such as a convolutional autoencoder model) for dimensionality reduction and reconstruction to obtain the background voiceprint feature of the advertising delivery scene. The encoder model maps the voiceprint sub-features of all input signal vectors to a low-dimensional latent space (such as a latent space with an encoding dimension of 16), and uses the reconstruction error in the latent space to determine which signal vector in the advertising delivery scene has the same voiceprint sub-feature as the background noise, and the signal vector with a high reconstruction error is regarded as speech noise or burst noise and is eliminated. The low-dimensional vector output after encoding all the voiceprint sub-features after noise elimination is used as the background voiceprint feature.
[0026] In some embodiments, reference Figure 2 This figure is a flowchart of an operation for determining an environmental perception state tensor according to some embodiments of the present application. In this application, cross-modal feature alignment is performed on the background voiceprint feature and the frequency domain feature of the light intensity in the light intensity data to obtain the environmental perception state tensor for the advertising delivery scene. This can be achieved by using the following steps: extracting frequency domain features of light intensity from the light intensity data; Performing time domain registration on the background voiceprint features and the frequency domain features to obtain the voiceprint-illumination dual-mode features of the advertising delivery scene; The voiceprint-illumination dual-mode features are mapped to a three-dimensional state space using a tensor fusion mechanism to form an environmental perception state tensor of the advertising delivery scenario.
[0027] In a specific implementation, the frequency domain features of the light intensity can be extracted from the light intensity data in the following manner: first, the light intensity data can be segmented according to a fixed time window length (such as 1 second) and the light intensity data segments under all time windows are subjected to short-time Fourier transform to calculate the average power value on each frequency component to obtain the power spectrum matrix of the light intensity; then, the power spectrum matrix is logarithmically transformed to form a eigenvector, and the eigenvector is used as the frequency domain feature of the light intensity; wherein the frequency domain feature is a eigenvector that characterizes the energy distribution of each frequency band of the light intensity signal, and the frequency domain feature can be used to capture the periodic changes of the ambient light, thereby providing accurate light time-frequency information for dynamic visual rendering.
[0028] In a specific implementation, the background voiceprint features and the frequency domain features are time-domain aligned to obtain the voiceprint-illumination dual-mode features of the advertising delivery scenario. This can be achieved in the following manner: first, the sampling frequencies of the background voiceprint features and the frequency domain features can be unified (e.g., the voiceprint features are grouped in a 10:1 ratio, averaged, and resampled to 10 Hz); then, a linear interpolation algorithm is used to align the alignment errors of the background voiceprint features and the frequency domain features on the time axis to ensure that the voiceprint vector and the illumination vector at each moment correspond one-to-one; then, the aligned background voiceprint features and the frequency domain features are merged in the input collaborative encoder (e.g., a multimodal neural network model) to generate a unified dual-mode feature vector (e.g., a "voiceprint 16-dimensional + illumination 50-dimensional" combined vector with a length of 66 dimensions), thereby obtaining a voiceprint-illumination dual-mode feature sequence of the advertising delivery scenario; wherein, the voiceprint-illumination dual-mode feature is a joint feature vector that characterizes the synergy between sound and illumination in the advertising delivery scenario, and the voiceprint-illumination dual-mode feature can be used to improve the accuracy of multimodal perception.
[0029] It should be noted that, in this application, the environmental perception state tensor is a high-dimensional data structure that characterizes the spatiotemporal correlation characteristics and dynamic change laws of sound and light in the advertising delivery scene. The environmental perception state tensor can provide multi-dimensional environmental state input for advertising rendering, so that the system can dynamically adjust the dynamic lighting matrix according to the coordinated changes of scene sound and light, enhance the spatiotemporal consistency between the advertising content and the real environment, improve the realism and environmental adaptability of the picture rendering, and ensure that the advertisement is naturally integrated into the complex scene and effectively conveys information; in specific implementation, the tensor fusion mechanism is used to map the voiceprint-light dual-mode features into a three-dimensional state space to form an environmental perception state tensor for the advertising delivery scene. This can be achieved in the following way: a multimodal tensor decomposition algorithm (such as canonical decomposition / parallel factor decomposition) can be used to map the aligned voiceprint-illumination dual-modal features into a three-dimensional tensor space to obtain the environmental perception state tensor of the advertising delivery scenario; wherein, the multimodal tensor decomposition algorithm reshapes the voiceprint-illumination feature matrix into a two-dimensional tensor (i.e., time × feature), and introduces a third dimension (such as spatial position) to construct a three-dimensional tensor (i.e., time × feature × space), and then decomposes the three-dimensional tensor into the sum of the outer products of three low-rank matrices through canonical decomposition / parallel factor decomposition, extracting cross-modal latent factors while retaining temporal dynamics and spatial structure information, thereby obtaining the environmental perception state tensor.
[0030] In step 103, a dynamic lighting matrix for rendering the target advertisement is generated by combining the spectral irradiance distribution in the environmental perception state tensor with the anisotropic highlight coefficients of each sub-region in the target advertisement screen.
[0031] In some embodiments, reference Figure 3This figure is an exemplary flow chart for determining a dynamic lighting matrix according to some embodiments of the present application. In the present application, the dynamic lighting matrix for rendering a target advertisement is generated by combining the spectral irradiance distribution in the environmental perception state tensor with the anisotropic highlight coefficients of each sub-region in the target advertisement image. The following steps can be used: In step 1031, the target advertising screen is divided into sub-regions to obtain multiple sub-regions; In step 1032, the anisotropic highlight coefficient of each sub-region in the target advertisement image is extracted; In step 1033, an environment matching factor of each sub-region is determined based on the spectral irradiance distribution in the environment perception state tensor; In step 1034 , the environment matching factors and anisotropic highlight coefficients of all sub-regions are mapped to the lighting rendering parameter space to generate a dynamic lighting matrix for rendering the target advertisement.
[0032] In specific implementation, the target advertising screen is segmented into sub-regions to obtain multiple sub-regions, which can be achieved in the following way: the target advertising screen can be input into a pre-trained deep segmentation network (such as a U-shaped network image segmentation model), and the instance segmentation algorithm is used to automatically identify and divide multiple sub-regions with similar texture features and output the pixel mask of each sub-region, thereby obtaining multiple sub-regions; wherein, the sub-region refers to an independent region with similar texture features divided from the advertising screen, and the sub-region can be used to specifically adjust the rendering parameters of different regions to improve the overall visual consistency.
[0033] In specific implementation, the extraction of the anisotropic highlight coefficient of each sub-region in the target advertising screen can be achieved in the following manner: for each sub-region, a bidirectional reflectance distribution function model (such as the Blinn-Phong model) can be used to fit the specular reflectance vector of the sub-region based on the physical rendering material of the sub-region, and the principal component analysis of the specular reflectance vector is performed on the specular reflectance vector to obtain the anisotropic highlight coefficient of the sub-region. The anisotropic highlight coefficient of each sub-region in the target advertising screen can be obtained through the above steps; wherein, the anisotropic highlight coefficient is a parameter that characterizes the change in specular reflection intensity of the material of the sub-region in the advertising screen under different incident angles. The anisotropic highlight coefficient can be used to accurately simulate the real lighting response of the regional surface and enhance the realism of the rendering.
[0034] In specific implementation, the environmental matching factor of each sub-region is determined based on the spectral irradiance distribution in the environmental perception state tensor, which can be achieved in the following manner, namely: the spectral irradiance distribution in the current time window can be extracted from the environmental perception state tensor through tensor slicing operation, and for each sub-region, the spectral irradiance vector of the sub-region is extracted from the spectral irradiance distribution, and the dot product of the spectral irradiance vector and the specular reflectivity vector of the sub-region is calculated to obtain the theoretical reflected light intensity of the sub-region under the current lighting environment, and the theoretical reflected light intensity is used as the environmental matching factor of the sub-region. The environmental matching factor of each sub-region can be obtained through the above steps; wherein, the environmental matching factor is a measure of the degree of adaptation of the sub-region under the current ambient lighting conditions, and the environmental matching factor can ensure that the rendering result maintains a natural transition in dynamic scenes; the spectral irradiance distribution is a distribution image that characterizes the radiation energy distribution density of the scene lighting.
[0035] It should be noted that in this application, the dynamic lighting matrix is a real-time rendering transformation matrix based on environmental perception and material characteristics. The dynamic lighting matrix can be used to accurately map the advertising image into a three-dimensional scene for lighting consistency rendering to achieve a dynamic rendering effect with consistent lighting and viewing angle; in specific implementation, the environmental matching factors and anisotropic highlight coefficients of all sub-areas are mapped to the lighting rendering parameter space, and the dynamic lighting matrix generated when rendering the target advertisement can be achieved in the following way, namely: the environmental matching factor and anisotropic highlight coefficient of each sub-area can be fused at the pixel level to construct a regional-level reflection modulation matrix, and then the reflection modulation matrix is mapped to the standard lighting rendering matrix format (such as 4×4) to obtain the dynamic lighting matrix when rendering the target advertisement.
[0036] In step 104, spatial depth data of the advertising delivery scene is obtained, and the audience position is identified based on the spatial depth data and the audience movement trajectory of the target advertising delivery scene to obtain the audience spatial domain where the target advertisement is played. Then, spatial audio redirection of the target advertisement is performed based on the audience spatial domain and the sound field energy distribution in the environmental perception state tensor to obtain the acoustic beam width angle when rendering the target advertisement.
[0037] It should be noted that, in this application, spatial depth data refers to point cloud data that reflects the three-dimensional distance information from each pixel point to the sensor in the advertising delivery scene. This spatial depth data can be used to reconstruct the three-dimensional structure of the advertising delivery scene and assist in accurately locating the audience position, thereby achieving accurate audio and video rendering and interaction; in specific implementation, obtaining the spatial depth data of the advertising delivery scene can be achieved in the following way, namely: in the advertising delivery scene, the point cloud data of the advertising delivery scene can be obtained through a structured light depth sensor (such as Intel RealSense D435) and the point cloud data can be used as the spatial depth data of the advertising delivery scene.
[0038] In some embodiments, the audience position is identified based on the spatial depth data and the audience movement trajectory in the target advertisement delivery scene, and the audience spatial domain where the target advertisement is played is obtained by the following steps: generating a three-dimensional point cloud map of the target advertising delivery scene based on the spatial depth data; Collect millimeter-wave radar signals of target advertising scenes and extract audience movement trajectories; Performing spatial fusion registration on the three-dimensional point cloud image and the audience's movement trajectory to obtain the spatial position distribution of the target audience; Sub-areas are divided based on the spatial position distribution to obtain the audience spatial domain for target advertisement broadcasting.
[0039] In specific implementation, generating a three-dimensional point cloud map of the target advertising delivery scene based on the spatial depth data can be achieved in the following manner, namely: the spatial depth data can be filtered using filtering tools in the point cloud library (such as voxel grid filtering) to remove noise and redundant points, thereby obtaining a dense three-dimensional point cloud map of objects and people in the advertising delivery scene, thereby obtaining a three-dimensional point cloud map of the target advertising delivery scene; wherein, the three-dimensional point cloud map is an image reflecting the three-dimensional distance information from each pixel point in the advertising delivery scene to the sensor, and the three-dimensional point cloud map can be used to reconstruct the real geometric structure of the advertising delivery scene, providing an accurate spatial reference for audience positioning and environmental understanding.
[0040] In a specific implementation, collecting millimeter-wave radar signals in the target advertising scene and then extracting audience movement trajectories can be achieved in the following manner: first, a millimeter-wave radar can be deployed in the advertising scene and continuous wave transmission can be enabled. The millimeter-wave radar hardware collects echo signals to obtain millimeter-wave radar signals and generates a range-Doppler matrix of the millimeter-wave radar signals. Then, using an existing signal detection algorithm (such as a constant false alarm rate detection algorithm), peak detection is performed on the range-Doppler matrix to extract a target echo list for each time frame. Multi-target tracking is performed on the target echo lists of all time frames, fusing distance and velocity information to obtain the real-time movement trajectory of each audience object in the advertising scene in two-dimensional polar coordinates. The collection of the real-time movement trajectories of all audience objects is then used as the audience movement trajectory. The audience movement trajectory is a position time series that represents the continuous movement path of the audience in the advertising scene. The audience movement trajectory can capture audience dynamics and accurately track their position change trends, providing a dynamic basis for targeted advertising rendering, allowing the advertising content to be flexibly adjusted as the audience moves, and enhancing the interactive experience.
[0041] In specific implementation, the three-dimensional point cloud map is spatially fused and aligned with the audience movement trajectory to obtain the spatial position distribution of the target audience, which can be achieved in the following manner, namely: first, the two-dimensional polar coordinate trajectory points in the audience movement trajectory can be converted into three-dimensional points through the coordinate transformation formula and mapped to the three-dimensional coordinate system of the structured light depth sensor; then, a clustering algorithm (such as a spatial clustering algorithm based on Euclidean distance) is used to identify the three-dimensional point clusters corresponding to the crowd in the three-dimensional point cloud map, and the trajectory points in the audience movement trajectory mapped to the three-dimensional coordinate system are paired with the center of the point cloud cluster for nearest neighbors, and the spatial position distribution of each audience object in the three-dimensional coordinate system after pairing is output; wherein, the spatial position distribution is a comprehensive data map reflecting the real-time position of all viewers in the advertising delivery scene and the spatial relationship between the three-dimensional scene. The spatial position distribution clearly marks the specific position of the audience in the space, providing an accurate position association basis for subsequent regional division and precise advertising delivery.
[0042] It should be noted that in this application, the audience space domain is a three-dimensional space area that includes the core audience group and is suitable for effective coverage of the target advertisement. The audience space domain can be used to define the target coverage range during advertisement playback and realize the precise interaction of spatial audio directional projection and visual content. In specific implementation, the sub-area division is performed based on the spatial position distribution, and the audience space domain for target advertisement playback can be obtained in the following manner, namely: the spatial position distribution can be projected onto the ground plane to obtain the two-dimensional coordinates of each audience on the horizontal plane and the audience position in the spatial position distribution can be partitioned on the ground plane using a clustering algorithm (such as a density-based spatial clustering algorithm) to obtain multiple sub-areas covering the audience. Then, all sub-areas are screened through the occlusion detection tool in the point cloud library (such as the ExtractIndices module) combined with the geometric structure of the advertising delivery scene in the three-dimensional point cloud map to eliminate invalid or occluded sub-areas, and the set of all sub-areas after screening is used as the audience space domain.
[0043] In some embodiments, spatial audio redirection of a target advertisement based on the audience spatial domain and the sound field energy distribution in the environmental perception state tensor to obtain an acoustic beamwidth angle when rendering the target advertisement can be achieved by using the following steps: Determining an auditory pointing vector for each audience member based on audience distribution and direction information in the audience space domain; Generate the audience area coverage angle of the target advertising scene through all auditory pointing vectors; Adjusting the frequency domain of the main audio channel according to the sound field energy distribution in the environmental perception state tensor to obtain an audio projection angle; An acoustic beamwidth angle for rendering a target advertisement is generated based on the audience area coverage angle and the audio projection angle.
[0044] In specific implementation, the auditory pointing vector of each audience member is determined based on the audience distribution and direction information in the audience space domain, which can be achieved in the following manner: for each audience sub-area in the audience space domain, the geometric center of the audience sub-area can be connected with the advertising sound source position coordinates and the unit vector between the two points is calculated on the horizontal plane to obtain the auditory pointing vector of the audience in the audience sub-area. The above steps can be used to obtain the auditory pointing vector of the audience in each audience sub-area in the audience space domain, thereby obtaining the auditory pointing vector of each audience member; wherein, the auditory pointing vector is a three-dimensional unit vector that represents the direction in which the audience expects to receive audio. The auditory pointing vector indicates the position direction from the sound source to the center of the audience sub-area where the audience is located, accurately captures the audience's auditory direction needs, and provides a core reference for audio directional transmission.
[0045] In specific implementation, the audience area coverage angle of the target advertising delivery scenario can be generated by all auditory pointing vectors in the following manner, namely: each auditory pointing vector can be projected onto the horizontal plane and the polar angle between each auditory pointing vector and the reference vector directly in front of the sound source can be calculated, and then the minimum and maximum values of all polar angles are taken as boundaries to obtain the horizontal angle interval covering all audiences in the target advertising delivery scenario, and the horizontal angle interval is used as the audience area coverage angle of the target advertising delivery scenario; wherein, the audience area coverage angle is an angular parameter that reflects the distribution range of the auditory needs of the audience group in the target advertising delivery scenario, and the audience area coverage angle can be used to define the horizontal range that the sound beam needs to cover, help determine the key areas that the audio needs to cover, and avoid invalid propagation.
[0046] In a specific implementation, the frequency domain of the audio main channel is adjusted according to the sound field energy distribution in the environmental perception state tensor, and the audio projection angle can be obtained in the following manner, namely: first, the power spectrum density distribution of the background voiceprint under the current time window can be extracted from the environmental perception state tensor through the tensor slicing operation, and the frequency domain noise in the power spectrum density distribution is weightedly analyzed (such as reducing the corresponding frequency band gain in the area with higher frequency band noise energy, and increasing the gain in the frequency band with lower energy to achieve spectrum balance) to obtain the sound field energy distribution; then, the existing beamforming technology (such as distributed adaptive beamforming algorithm) is used to search in the horizontal plane through the sound field energy distribution to make the signal-to-noise ratio equalized. The projection direction with the highest ratio is selected, and the projection direction is used as the audio projection angle; wherein, the audio projection angle is the best propagation direction angle that characterizes the maximum avoidance of environmental noise interference. The audio projection angle combined with the acoustic beam width angle can dynamically optimize the audio propagation direction, reduce the interference of background noise on advertising audio, enhance audio clarity and penetration, and ensure that advertising sounds can still be accurately conveyed in complex environments; the sound field energy distribution is the intensity distribution feature that characterizes the frequency composition and energy distribution state of background noise in the advertising delivery scene. The sound field energy distribution can provide an environmental noise feature basis for spatial audio redirection, and avoid high-energy noise frequency bands by adjusting the frequency domain of the main audio channel.
[0047] It should be noted that, in the present application, the acoustic beam width angle is the horizontal angle range that the directional audio emitted from the sound source should cover. The acoustic beam width angle can balance the audience coverage range and the audio propagation accuracy, and provide a quantitative basis for advertising audio rendering, so that the advertising sound can not only cover the target audience group, but also reduce the sound diffusion loss, and improve the overall effect of the synchronous rendering of advertising sound and light; in specific implementation, the acoustic beam width angle for rendering the target advertisement based on the audience area coverage angle and the audio projection angle can be achieved in the following way, that is: the audience area coverage angle and the audio projection angle can be combined to calculate the acoustic beam width angle for rendering the target advertisement, for example: with the audio projection angle as the center, the maximum value of the deviation between the left and right boundaries of the coverage angle interval and the center angle is taken as the half width, and then the half width is multiplied by two to obtain the full width angle as the acoustic beam width angle. In other embodiments, other methods can also be used to combine the audience area coverage angle with the audio projection angle, which is not specifically limited here.
[0048] In step 105, the advertisement content of the target advertisement is dynamically rendered with sound and light synchronization according to the dynamic illumination matrix and the acoustic beamwidth angle.
[0049] In some embodiments, performing dynamic acoustic and optical synchronous rendering of the advertisement content of the target advertisement according to the dynamic illumination matrix and the acoustic beamwidth angle may be implemented by the following steps: Adjusting camera parameters in a rendering engine pipeline based on the dynamic lighting matrix to achieve dynamic lighting rendering; adjusting the amplitude and phase of the multi-channel audio according to the acoustic beamwidth angle; A time synchronization lock mechanism is constructed to control the synchronization error between the visual rendering and audio rendering of the target advertising content, and then the advertising content of the target advertisement is dynamically rendered in sound and light synchronization according to the synchronization error.
[0050] In specific implementation, the camera parameters in the rendering engine pipeline are adjusted based on the dynamic lighting matrix, and dynamic lighting rendering can be achieved in the following way: the dynamic lighting matrix can be passed as an input parameter to the lighting rendering component in the universal rendering pipeline of the rendering engine (such as the Unity engine), and the programmable rendering pipeline characteristics in the pipeline are used to read the matrix data through code and dynamically adjust the exposure compensation parameters, white balance coefficient, gamma correction factor and other lighting rendering parameters to perform dynamic lighting rendering on the advertising screen.
[0051] In specific implementation, adjusting the amplitude and phase of multi-channel audio according to the acoustic beamwidth angle can be achieved in the following manner, namely: the acoustic beamwidth angle can be input into professional audio middleware (such as the FMOD audio engine) as a half-angle parameter of the directional sound source in the audio rendering pipeline to adjust the amplitude and phase of the multi-channel audio; wherein, the multi-channel beamforming algorithm in the audio engine divides the stereo or multi-channel audio track of the target advertisement into several sub-sound sources, and dynamically allocates sound source weights according to the acoustic beamwidth angle to adjust the amplitude and phase of each channel. During operation, the audio engine will calculate and apply these amplitude and phase offsets in real time, concentrate the sound energy within the coverage angle range of the audience area, and suppress leakage in other directions, so as to achieve precise directional projection of the sound field.
[0052] In specific implementation, a time synchronization lock mechanism is constructed to control the synchronization error between the visual rendering and audio rendering of the target advertising content. Then, dynamic rendering of the target advertising content in synchronization with sound and light can be performed based on the synchronization error. This can be achieved in the following way: based on the high-precision timer function in the real-time operating system (such as VxWorks RTX), a global timestamp signal can be created as the synchronization benchmark for visual and audio rendering. In the rendering pipeline, a timestamp is generated for each visual frame and bound to the timestamp of the audio sampling point. A double-buffer queue mechanism is then used to store the visual frames and audio data blocks to be rendered separately. The consistency of the timestamps of the two is ensured through a hardware timestamp synchronization interface (such as the IEEE 1588 Precision Time Protocol). When the synchronization error is detected to exceed a preset error threshold (such as 5ms), the rendering speed is dynamically adjusted using a feedback control algorithm. For example, if the audio is ahead, the audio playback thread is delayed; if the visual is ahead, the visual rendering is accelerated or some non-key frames are skipped, ultimately achieving high-precision synchronization of the sound and light rendering of the advertising content.
[0053] In addition, in another aspect of the present application, in some embodiments, the present application provides a visual communication advertising design system, the system includes an advertising rendering unit, reference Figure 4 , which is a schematic diagram of the structure of an advertisement rendering unit according to some embodiments of the present application. The advertisement rendering unit 400 includes: a collection module 401, a processing module 402, and an execution module 403, which are described as follows: The acquisition module 401 in this application is mainly used to synchronously collect light intensity data and environmental voiceprint data of the advertising scene through the Internet of Things sensor network; Processing module 402, in this application, is mainly used to extract background voiceprint features of the advertising delivery scene from the environmental voiceprint data, and then perform cross-modal feature alignment on the background voiceprint features and the frequency domain features of the light intensity in the light intensity data to obtain an environmental perception state tensor of the advertising delivery scene; It should be noted that the processing module 402 in the present application is also used to generate a dynamic lighting matrix for rendering the target advertisement by combining the spectral irradiance distribution in the environmental perception state tensor with the anisotropic highlight coefficients of each sub-region in the target advertisement screen; In addition, it should be noted that the processing module 402 in the present application is also used to obtain spatial depth data of the advertising delivery scene, identify the audience position according to the spatial depth data and the audience movement trajectory of the target advertising delivery scene, obtain the audience spatial domain where the target advertisement is played, and then perform spatial audio redirection on the target advertisement based on the audience spatial domain and the sound field energy distribution in the environmental perception state tensor to obtain the acoustic beamwidth angle when rendering the target advertisement; The execution module 403 in this application is mainly used to perform sound and light synchronous dynamic rendering of the advertisement content of the target advertisement according to the dynamic lighting matrix and the acoustic beam width angle.
[0054] Each module in the visual communication advertising design system described above may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a computer device memory in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0055] In addition, in one embodiment, the present application provides a computer device, which may be a server, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data of an advertising rendering method based on sound and light coordination and environmental perception. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, an advertising rendering method based on sound and light coordination and environmental perception is implemented.
[0056] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0057] In one embodiment, a computer device is also provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the above-mentioned embodiment of the advertising rendering method based on sound and light coordination and environmental perception are implemented.
[0058] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned embodiment of the advertisement rendering method based on sound-light collaboration and environmental perception are implemented.
[0059] In one embodiment, a computer program product or program is provided. The computer program product or program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of the aforementioned embodiment of the method for rendering advertisements based on audio-visual collaboration and environmental perception.
[0060] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0061] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0062] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. An advertisement rendering method based on sound and light synergy and environmental perception, used for dynamically rendering designed advertisements in a visual communication advertisement design system, characterized in that: The method comprises the following steps: Synchronously collect light intensity data and environmental voiceprint data of advertising scenes through the IoT sensor network; Extracting background voiceprint features of the advertising delivery scene from the environmental voiceprint data, and then performing cross-modal feature alignment on the background voiceprint features and the frequency domain features of light intensity in the light intensity data to obtain an environmental perception state tensor of the advertising delivery scene; Generate a dynamic lighting matrix for rendering the target advertisement by combining the spectral irradiance distribution in the environmental perception state tensor with the anisotropic highlight coefficients of each sub-region in the target advertisement screen; Acquiring spatial depth data of an advertisement delivery scene, identifying the audience's position based on the spatial depth data and the audience's movement trajectory in the target advertisement delivery scene, obtaining the audience spatial domain where the target advertisement is played, and then performing spatial audio redirection on the target advertisement based on the audience spatial domain and the sound field energy distribution in the environmental perception state tensor to obtain the acoustic beamwidth angle when rendering the target advertisement; The advertisement content of the target advertisement is dynamically rendered in a sound and light synchronous manner according to the dynamic illumination matrix and the acoustic beam width angle.
2. The method according to claim 1, wherein Extracting background voiceprint features of the advertising delivery scene from the environmental voiceprint data specifically includes: Performing a short-time Fourier transform on the environmental voiceprint data to obtain a sound spectrum of the environmental voiceprint data; Mapping the environmental voiceprint data into a sound field static texture based on the sound spectrogram to obtain a voiceprint sub-feature of each signal vector in the environmental voiceprint data; All voiceprint sub-features are reconstructed by dimensionality reduction to obtain the background voiceprint features of the advertising scene.
3. The method according to claim 1, wherein Performing cross-modal feature alignment on the background voiceprint feature and the frequency domain feature of the light intensity in the light intensity data to obtain the environment perception state tensor of the advertising delivery scene specifically includes: extracting frequency domain features of light intensity from the light intensity data; Performing time domain registration on the background voiceprint features and the frequency domain features to obtain the voiceprint-illumination dual-mode features of the advertising delivery scene; The voiceprint-illumination dual-mode features are mapped to a three-dimensional state space using a tensor fusion mechanism to form an environmental perception state tensor of the advertising delivery scenario.
4. The method according to claim 1, wherein Generating a dynamic lighting matrix for rendering the target advertisement by combining the spectral irradiance distribution in the environment perception state tensor with the anisotropic highlight coefficients of each sub-region in the target advertisement screen specifically includes: Divide the target advertising screen into sub-regions to obtain multiple sub-regions; Extracting the anisotropic highlight coefficient of each sub-region in the target advertising image; determining an environment matching factor for each sub-region based on the spectral irradiance distribution in the environment perception state tensor; The environment matching factors and anisotropic highlight coefficients of all sub-areas are mapped to the lighting rendering parameter space to generate a dynamic lighting matrix for rendering the target advertisement.
5. The method according to claim 1, wherein Identifying the audience's position based on the spatial depth data and the audience's movement trajectory in the target advertisement delivery scene, and obtaining the audience spatial domain for the target advertisement playback specifically includes: generating a three-dimensional point cloud map of the target advertising delivery scene based on the spatial depth data; Collect millimeter-wave radar signals of target advertising scenes and extract audience movement trajectories; Performing spatial fusion registration on the three-dimensional point cloud image and the audience's movement trajectory to obtain the spatial position distribution of the target audience; Sub-areas are divided based on the spatial position distribution to obtain the audience spatial domain for target advertisement broadcasting.
6. The method according to claim 1, wherein Performing spatial audio redirection on a target advertisement based on the audience spatial domain and the sound field energy distribution in the environmental perception state tensor to obtain an acoustic beamwidth angle when rendering the target advertisement specifically includes: Determining an auditory pointing vector for each audience member based on audience distribution and direction information in the audience space domain; Generate the audience area coverage angle of the target advertising scene through all auditory pointing vectors; Adjusting the frequency domain of the main audio channel according to the sound field energy distribution in the environmental perception state tensor to obtain an audio projection angle; An acoustic beamwidth angle for rendering a target advertisement is generated based on the audience area coverage angle and the audio projection angle.
7. The method according to claim 1, wherein Performing acoustic and optical synchronous dynamic rendering of the advertisement content of the target advertisement according to the dynamic illumination matrix and the acoustic beam width angle specifically includes: Adjusting camera parameters in a rendering engine pipeline based on the dynamic lighting matrix to achieve dynamic lighting rendering; adjusting the amplitude and phase of the multi-channel audio according to the acoustic beamwidth angle; A time synchronization lock mechanism is constructed to control the synchronization error between the visual rendering and audio rendering of the target advertising content, and then the advertising content of the target advertisement is dynamically rendered in sound and light synchronization according to the synchronization error.
8. A visual communication advertising design system, comprising an advertising rendering unit, characterized in that: The advertisement rendering unit includes: The acquisition module is used to synchronously collect light intensity data and environmental voiceprint data of the advertising scene through the IoT sensor network; a processing module configured to extract background voiceprint features of the advertising delivery scene from the environmental voiceprint data, and then perform cross-modal feature alignment on the background voiceprint features and the frequency domain features of light intensity in the light intensity data to obtain an environmental perception state tensor of the advertising delivery scene; The processing module is configured to generate a dynamic lighting matrix for rendering the target advertisement by combining the spectral irradiance distribution in the environmental perception state tensor with the anisotropic highlight coefficients of each sub-region in the target advertisement screen; The processing module is configured to obtain spatial depth data of an advertisement delivery scene, identify an audience position based on the spatial depth data and an audience movement trajectory of the target advertisement delivery scene, obtain an audience spatial domain for the target advertisement, and then perform spatial audio redirection on the target advertisement based on the audience spatial domain and the sound field energy distribution in the environmental perception state tensor to obtain an acoustic beamwidth angle when rendering the target advertisement; An execution module is configured to perform sound and light synchronous dynamic rendering on the advertisement content of the target advertisement according to the dynamic illumination matrix and the acoustic beam width angle.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the advertisement rendering method based on sound-light coordination and environmental perception according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the advertisement rendering method based on sound-light coordination and environmental perception are implemented as described in any one of claims 1 to 7.
Citation Information
Cited By
Listening training system for English teaching
CN120954287A
A listening training system for English teaching
CN120954287B
Urban landmark feature extraction method based on AI model
CN121213942A
Advertisement push picture generation method and system based on style migration
CN121353454A
Advertisement pushing picture generation method and system based on style transfer
CN121353454B