Environment atmosphere control method and device based on data analysis, equipment and medium
Patent Information
- Application Number
- CN202510856571.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-06-25
AI Technical Summary
[0003]有鉴于此,本申请的目的在于提供一种基于数据分析的环境氛围控制方法和装置、设备及介质,以改善现有技术中存在的环境氛围控制的效果相对不佳的问题
[0014]本申请提供的基于数据分析的环境氛围控制方法和装置、设备及介质,首先,获取针对目标歌曲预先配置的多个候选氛围控制方案;其次,获取目标环境中的目标音频数据;然后,获取目标环境中的目标图像数据;之后,基于目标图像数据具有的语义信息,对目标音频数据具有的语义信息进行语义强化操作,形成强化音频语义特征;最后,基于强化音频语义特征,在多个候选氛围控制方案中,确定出与目标环境相匹配的目标氛围控制方案。基于上述内容,一方面,由于会在多个候选氛围控制方案中确定目标氛围控制方案,使得目标氛围控制方案的确定范围相对较小,因此,可以保障确定目标氛围控制方案的精度,另一方面,由于会基于具有目标用户的姿态信息的目标图像数据,对目标音频数据进行语义强化,使得形成强化音频语义特征对环境氛围的表征准确度和丰富度都能够得到提升,因此,可以保障确定的目标氛围控制方案的可靠度,使得能够对环境氛围进行可靠地控制,从而改善现有技术中存在的环境氛围控制的效果相对不佳的问题。
Smart Images

Figure CN120730586B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to an environmental atmosphere control method, apparatus, device, and medium based on data analysis. Background Technology
[0002] In today's digital age, ambient atmosphere control technology has been widely used in entertainment, home, office, and public spaces. However, existing technologies still have some limitations in ambient atmosphere control. Traditional methods either rely on fixed atmosphere control schemes or depend on the rhythm of music to adjust the atmosphere. This simplistic approach often fails to fully reflect the true state of the environment, resulting in poor control effectiveness. Summary of the Invention
[0003] In view of this, the purpose of this application is to provide an environmental atmosphere control method, apparatus, device and medium based on data analysis, so as to improve the problem of relatively poor environmental atmosphere control effect in the prior art.
[0004] To achieve the above objectives, this application adopts the following technical solution: An environmental atmosphere control method based on data analysis includes: Obtain multiple candidate atmosphere control schemes pre-configured for a target song, wherein the target song is the song currently being sung by the target user in the target environment, and the candidate atmosphere control schemes include at least a scheme for controlling the lighting in the target environment; Acquire target audio data in the target environment, wherein the target audio data includes the sound of the target user singing the target song and background sound; Acquire target image data in the target environment, wherein the target image data is used to reflect at least the posture information of the target user; Based on the semantic information of the target image data, semantic enhancement operations are performed on the semantic information of the target audio data to form enhanced audio semantic features. In the enhanced audio semantic features, the semantic information of the target audio data is the main component, and the semantic information of the target image data is the auxiliary component. Based on the enhanced audio semantic features, a target atmosphere control scheme that matches the target environment is determined from among the multiple candidate atmosphere control schemes. The target atmosphere control scheme is used to control the lighting in the target environment.
[0005] In a preferred embodiment of this application, in the aforementioned environmental atmosphere control method based on data analysis, the step of performing semantic enhancement operations on the semantic information of the target audio data based on the semantic information of the target image data to form enhanced audio semantic features includes: Semantic extraction is performed on the target image data to form semantic features of the target image; Semantic extraction is performed on the target audio data to form target audio semantic features; Based on the target enhancement matrix and the target image semantic features, semantic enhancement operations are performed on the target audio semantic features to form enhanced audio semantic features. The target enhancement matrix is formed by learning from sample audio data, sample image data and sample atmosphere control scheme, and the initial target enhancement matrix is a random matrix.
[0006] In a preferred embodiment of this application, in the aforementioned data analysis-based environmental atmosphere control method, the step of performing semantic enhancement operations on the target audio semantic features based on the target enhancement matrix and the target image semantic features to form enhanced audio semantic features includes: Based on the feature size of the semantic features of the target image and the feature size of the semantic features of the target audio, a first enhancement sub-matrix whose feature size matches the feature size of the semantic features of the target image is extracted from the target enhancement matrix, and a second enhancement sub-matrix whose feature size matches the feature size of the semantic features of the target audio is extracted. Based on the first enhancement sub-matrix, the semantic features of the target image are linearly mapped to form target image mapping features; Based on the second enhancement sub-matrix, the target audio semantic features are linearly mapped to form target audio mapping features; Based on the target image mapping features, cross-attention processing and residual connections are performed on the target audio mapping features to form enhanced audio semantic features.
[0007] In a preferred embodiment of this application, in the aforementioned data analysis-based environmental atmosphere control method, the step of performing semantic enhancement operations on the target audio semantic features based on the target enhancement matrix and the target image semantic features to form enhanced audio semantic features includes: The semantic features of the target image and the semantic features of the target audio are compressed to form target image compressed features and target audio compressed features of multiple feature sizes; In the first stage, based on the first enhancement matrix included in the target enhancement matrix and the target image compression feature with the largest feature size, a semantic enhancement operation is performed on the target audio compression feature with the largest feature size to form the enhancement feature of the first stage. There is a one-to-one correspondence between the multiple feature sizes and the number of enhancement stages. In each of the second and subsequent stages, the enhancement features of the previous stage are connected to the target audio compression features of the corresponding feature size, and semantic enhancement is performed on the connection result based on the enhancement matrix of the current stage and the target image compression features of the corresponding feature size included in the target enhancement matrix, to form the enhancement features of the current stage. Based on the enhancement features of the last stage, the enhanced audio semantic features are determined.
[0008] In a preferred embodiment of this application, in the aforementioned data analysis-based environmental atmosphere control method, the steps of connecting the enhancement features of the previous stage to the target audio compression features of the corresponding feature size in each of the second and subsequent stages, and performing semantic enhancement operations on the connection results based on the enhancement matrix of the current stage and the target image compression features of the corresponding feature size included in the target enhancement matrix, to form the enhancement features of the current stage, include: In each of the second and subsequent stages, the enhancement features from the previous stage and the target audio compression features of the corresponding feature size are concatenated, and the concatenation result is pooled and compressed to form the connection result of the current stage. Based on the target enhancement matrix, which includes the enhancement matrix of the current stage and the target image compression features of the corresponding feature size, semantic enhancement operations are performed on the connection results of the current stage to form the enhancement features of the current stage.
[0009] In a preferred embodiment of this application, in the above-described environmental atmosphere control method based on data analysis, the step of performing semantic extraction on the target image data to form semantic features of the target image includes: The target image data is frequency domain converted to form first spectrum data; Multiple first mask matrices are used to mask the first spectrum data to form multiple first spectrum mask data. The multiple first mask matrices are formed by learning from sample audio data, sample image data and sample atmosphere control scheme. Attention features are mined for each of the first spectral mask data to form the first spectral attention feature corresponding to each of the first spectral mask data; Multiple first-spectrum attention features are fused to form target image semantic features.
[0010] In a preferred embodiment of this application, in the aforementioned data analysis-based environmental atmosphere control method, the step of performing semantic extraction on the target audio data to form target audio semantic features includes: The target audio data is frequency domain converted to form second spectrum data; Multiple second mask matrices are used to mask the second spectral data to form multiple second spectral mask data. The multiple second mask matrices are formed by learning from sample audio data, sample image data and sample atmosphere control scheme. Attention features are mined for each of the second spectral mask data to form the second spectral attention features corresponding to each of the second spectral mask data; Multiple second-spectrum attention features are fused to form target audio semantic features.
[0011] This application also provides an environmental atmosphere control device based on data analysis, including: The control scheme acquisition module is used to acquire multiple candidate atmosphere control schemes pre-configured for a target song, wherein the target song is the song currently being sung by the target user in the target environment, and the candidate atmosphere control schemes include at least a scheme for controlling the lights in the target environment. An audio data acquisition module is used to acquire target audio data in the target environment, wherein the target audio data includes the sound of the target user singing the target song and background sound; An image data acquisition module is used to acquire target image data in the target environment, wherein the target image data is used to reflect at least the posture information of the target user; The semantic enhancement module is used to perform semantic enhancement operations on the semantic information of the target audio data based on the semantic information of the target image data, forming enhanced audio semantic features. In the enhanced audio semantic features, the semantic information of the target audio data is the main component, and the semantic information of the target image data is the auxiliary component. The control scheme determination module is used to determine, based on the enhanced audio semantic features, a target atmosphere control scheme that matches the target environment from among the multiple candidate atmosphere control schemes, wherein the target atmosphere control scheme is used to control the lighting in the target environment.
[0012] Based on the above, this application also provides an electronic device, including: Memory, used to store computer programs; A processor connected to the memory is used to execute the computer program stored in the memory to implement the above-described environmental atmosphere control method based on data analysis.
[0013] Based on the above, this application also provides a computer-readable storage medium storing a computer program that, when executed, performs the various steps of the above-described data analysis-based environmental atmosphere control method.
[0014] The environmental atmosphere control method, apparatus, device, and medium based on data analysis provided in this application first acquire multiple candidate atmosphere control schemes pre-configured for a target song; second, acquire target audio data in the target environment; then, acquire target image data in the target environment; subsequently, based on the semantic information of the target image data, perform semantic enhancement operations on the semantic information of the target audio data to form enhanced audio semantic features; finally, based on the enhanced audio semantic features, determine the target atmosphere control scheme that matches the target environment from among the multiple candidate atmosphere control schemes. Based on the above, on the one hand, since the target atmosphere control scheme is determined from multiple candidate atmosphere control schemes, the range of determination for the target atmosphere control scheme is relatively small, thus ensuring the accuracy of the determined target atmosphere control scheme. On the other hand, since semantic enhancement is performed on the target audio data based on the target image data containing the target user's posture information, the accuracy and richness of the representation of the environmental atmosphere by the formed enhanced audio semantic features are improved, thus ensuring the reliability of the determined target atmosphere control scheme and enabling reliable control of the environmental atmosphere, thereby improving the relatively poor environmental atmosphere control effect in existing technologies. Attached Figure Description
[0015] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings.
[0016] Figure 1 A structural block diagram of an electronic device provided in an embodiment of this application.
[0017] Figure 2 This is a flowchart illustrating the environmental atmosphere control method based on data analysis provided in an embodiment of this application.
[0018] Figure 3 This is a schematic diagram illustrating the semantic extraction operation of image data provided in an embodiment of this application.
[0019] Figure 4 This is a schematic diagram illustrating the semantic extraction operation of audio data provided in an embodiment of this application.
[0020] Figure 5This is a schematic diagram illustrating the semantic enhancement operation provided in an embodiment of this application.
[0021] Figure 6 A block diagram of an environmental atmosphere control device based on data analysis provided in an embodiment of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0023] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0024] like Figure 1 As shown in the illustration, this application provides an electronic device. The electronic device may include a memory, a processor, and an environmental atmosphere control device based on data analysis.
[0025] Specifically, the memory and the processor are electrically connected directly or indirectly to enable data transmission or interaction. For example, the memory and the processor can be electrically connected via one or more communication buses or signal lines. The data analysis-based environmental atmosphere control device includes at least one software functional module stored in the memory in the form of software or firmware. The processor is used to execute executable computer programs stored in the memory, such as the software functional modules and computer programs included in the data analysis-based environmental atmosphere control device, to implement the data analysis-based environmental atmosphere control method provided in this application embodiment.
[0026] Optionally, the memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0027] Furthermore, the processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a system on chip (SoC), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0028] Understandable. Figure 1 The structure shown is for illustrative purposes only; the electronic device may also include components that are more advanced than those shown. Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown may include, for example, a communication unit for interacting with other devices (such as image sensors, audio acquisition units, controllers, etc.) to acquire the required images, audio, or output ambient control schemes.
[0029] Combination Figure 2 This application also provides a data analysis-based environmental atmosphere control method applicable to the aforementioned electronic device. The method steps defined in the process of the data analysis-based environmental atmosphere control method can be implemented by the electronic device.
[0030] The following will be about Figure 2 The specific process shown will be explained in detail.
[0031] Step S110: Obtain multiple candidate atmosphere control schemes pre-configured for the target song.
[0032] In this embodiment of the invention, the electronic device can acquire multiple candidate atmosphere control schemes pre-configured for the target song. For example, the multiple candidate atmosphere control schemes can be selected by the administrator from all atmosphere control schemes, or they can be generated directly. The target song is the song currently being sung by the target user in the target environment. The candidate atmosphere control schemes at least include schemes for controlling the lighting in the target environment. Exemplarily, in other embodiments, other devices can also be controlled, such as temperature control devices.
[0033] Step S120: Obtain target audio data in the target environment.
[0034] In this embodiment of the invention, the electronic device can acquire target audio data in the target environment. The target audio data includes the sound of the target user singing the target song and background sounds, such as sound captured by a microphone.
[0035] Step S130: Obtain target image data in the target environment.
[0036] In this embodiment of the invention, the electronic device can acquire target image data in the target environment. The target image data at least reflects the posture information of the target user; for example, it is generated by capturing images of the target user using an image acquisition device.
[0037] Step S140: Based on the semantic information of the target image data, perform semantic enhancement operations on the semantic information of the target audio data to form enhanced audio semantic features.
[0038] In this embodiment of the invention, after acquiring the target audio data and the target image data, the electronic device can perform semantic enhancement operations on the semantic information of the target audio data based on the semantic information of the target image data, forming enhanced audio semantic features. Specifically, in the enhanced audio semantic features, the semantic information of the target audio data is primary, while the semantic information of the target image data is secondary. That is, the target image data is used as an auxiliary tool to mine the semantic information in the target audio data. Thus, the enhanced audio semantic features can carry not only the semantic information of the target user's voice singing the target song and the background sound, but also the target user's posture information.
[0039] Step S150: Based on the enhanced audio semantic features, determine the target atmosphere control scheme that matches the target environment from among the multiple candidate atmosphere control schemes.
[0040] In this embodiment of the invention, after forming the enhanced audio semantic features, the electronic device can determine a target atmosphere control scheme that matches the target environment from among the multiple candidate atmosphere control schemes based on the enhanced audio semantic features. The target atmosphere control scheme is used to control the lighting in the target environment. For example, semantic features can be extracted from each candidate atmosphere control scheme (using any existing method for extracting semantic features from text, such as word embedding), to obtain the scheme semantic features corresponding to each candidate atmosphere control scheme. Then, the feature similarity (such as cosine similarity) between the enhanced audio semantic features and each scheme semantic feature can be calculated. Finally, the candidate atmosphere control scheme corresponding to the scheme semantic feature with the maximum feature similarity can be determined as the target atmosphere control scheme that matches the target environment.
[0041] Based on the above, on the one hand, since the target atmosphere control scheme is determined from multiple candidate atmosphere control schemes, the range of determination of the target atmosphere control scheme is relatively small. Therefore, the accuracy of determining the target atmosphere control scheme can be guaranteed. On the other hand, since semantic enhancement is performed on the target audio data based on the target image data with the target user's posture information, the accuracy and richness of the representation of the environmental atmosphere by the enhanced audio semantic features can be improved. Therefore, the reliability of the determined target atmosphere control scheme can be guaranteed, enabling reliable control of the environmental atmosphere, thereby improving the problem of relatively poor environmental atmosphere control in the existing technology.
[0042] Regarding the above-mentioned environmental atmosphere control method based on data analysis, it should be noted that the specific method of semantic enhancement operation on the semantic information of the target audio data is not limited and can be selected according to actual needs.
[0043] For example, in an alternative implementation, cross-attention processing can be performed on the semantic information of the target audio data (i.e., the corresponding semantic features obtained by semantic extraction) based on the semantic information of the target image data (i.e., the corresponding semantic features obtained by semantic extraction) to obtain enhanced audio semantic features.
[0044] For example, in another alternative implementation, the semantic information of the target image data (i.e., the corresponding semantic features obtained by semantic extraction) can be linearly mapped to form a gating feature. Then, the semantic information of the target audio data (i.e., the corresponding semantic features obtained by semantic extraction) can be gated and filtered (i.e., multiplied bitwise) based on the gating feature to obtain enhanced audio semantic features.
[0045] For example, in another alternative implementation, in order to improve the accuracy of semantic enhancement operations, the above step S140 may further include steps S141, S142 and S143, the specific contents of each step are as follows.
[0046] Step S141: Perform semantic extraction on the target image data to form target image semantic features.
[0047] In this embodiment, semantic extraction is performed on the target image data to form target image semantic features. That is, latent semantic information, such as pose-related semantic information, can be extracted from the target image data. Furthermore, in this embodiment, each feature and semantic feature can be represented as a vector.
[0048] Step S142: Perform semantic extraction on the target audio data to form target audio semantic features.
[0049] In this embodiment of the application, semantic extraction operations can be performed on the target audio data to form target audio semantic features. That is, potential semantic information in the target audio data, such as semantic information related to song sounds and environmental sounds, can be extracted.
[0050] Step S143: Based on the target enhancement matrix and the target image semantic features, perform semantic enhancement operation on the target audio semantic features to form enhanced audio semantic features.
[0051] In this embodiment, after forming the target image semantic features and the target audio semantic features, semantic enhancement operations can be performed on the target audio semantic features based on the target enhancement matrix and the target image semantic features to form enhanced audio semantic features. The target enhancement matrix is learned from sample audio data, sample image data, and sample atmosphere control schemes (i.e., labels), and the initial target enhancement matrix is a random matrix. Thus, the knowledge learned from the samples (i.e., the target enhancement matrix) can be fully utilized during the semantic enhancement operation, ensuring the reliability of the semantic enhancement operation.
[0052] It is understood that the specific method of performing semantic extraction on the target image data in step S141 above is not limited. For example, in an alternative implementation, in order to perform high-precision extraction on different semantic information in the semantic extraction operation, thereby ensuring the accuracy of the semantic representation of the obtained target image semantic features, step S141 above may further include steps S141a, S141b, S141c, and S141d. The specific contents of each step are as follows (in conjunction with...). Figure 3 (As shown).
[0053] Step S141a: The target image data is frequency domain converted to form first spectrum data.
[0054] In this embodiment, the target image data can be frequency domain transformed to form first spectral data. That is, the spatial domain target image data is transformed to the frequency domain, thus obtaining the corresponding spectrogram, i.e., the first spectral data. The specific method for performing the frequency domain transformation can be Fourier transform, and the specific processing procedure can refer to relevant prior art.
[0055] Step S141b: The first spectrum data is masked by multiple first mask matrices to form multiple first spectrum mask data.
[0056] In this embodiment, after forming the first spectral data, multiple first mask matrices can be used to mask the first spectral data, forming multiple first spectral mask data. These multiple first mask matrices are learned from sample audio data, sample image data, and sample atmosphere control schemes. For example, the first spectral data is masked using a first first mask matrix to form a first first spectral mask data; a second first mask matrix is used to mask the first spectral data to form a second first spectral mask data; and a third first mask matrix is used to mask the first spectral data to form a third first spectral mask data. Furthermore, it should be noted that each parameter in the first mask matrix is either 0 or 1, allowing the first mask matrix and the first spectral data to be multiplied bitwise to obtain the corresponding first spectral mask data. Based on this, multiple first mask matrices can mask different components in the first spectral data, enabling separate attention to different components and improving the accuracy of semantic extraction.
[0057] Step S141c: Mining attention features for each of the first spectral mask data to form the first spectral attention feature corresponding to each of the first spectral mask data.
[0058] In this embodiment of the application, after forming the first spectral mask data, attention features can be mined for each of the first spectral mask data to form a first spectral attention feature corresponding to each of the first spectral mask data. For example, for the first first spectral mask data, convolution processing can be performed on the first spectral mask data to obtain the corresponding convolution feature. Then, self-attention processing can be performed on the convolution feature to obtain the first spectral attention feature corresponding to the first first spectral mask data.
[0059] Step S141d: Fuse multiple first spectral attention features to form target image semantic features.
[0060] In this embodiment of the application, after forming the first spectral attention feature, multiple first spectral attention features can be fused to form the target image semantic feature; for example, multiple first spectral attention features can be added or averaged to form the target image semantic feature, or multiple first spectral attention features can be concatenated, and then the concatenated features can be sequentially processed by convolution, pooling and fully connected to form the target image semantic feature.
[0061] It is understood that the specific method of performing semantic extraction on the target audio data in step S142 above is not limited. For example, in an alternative embodiment, in order to perform high-precision extraction on different semantic information separately in the semantic extraction operation, thereby ensuring the accuracy of the semantic representation of the obtained target audio semantic features, step S142 above may further include steps S142a, S142b, S142c, and S142d. The specific contents of each step are as follows (in conjunction with...). Figure 4 (As shown).
[0062] Step S142a: The target audio data is frequency domain converted to form second spectrum data.
[0063] In this embodiment, the target audio data can be frequency domain transformed to form second spectral data. That is, the time-domain target audio data is transformed into the frequency domain, thus obtaining the corresponding spectrogram, i.e., the second spectral data. The specific method for performing the frequency domain transformation can be Fourier transform, and the specific processing procedure can refer to relevant prior art.
[0064] Step S142b: The second spectral data is masked using multiple second mask matrices to form multiple second spectral mask data.
[0065] In this embodiment, after forming the second spectral data, multiple second mask matrices can be used to mask the second spectral data, forming multiple second spectral mask data. These multiple second mask matrices are learned from sample audio data, sample image data, and sample atmosphere control schemes. For example, a first second mask matrix is used to mask the second spectral data to form a first second spectral mask data; a second second mask matrix is used to mask the second spectral data to form a second second spectral mask data; and a third second mask matrix is used to mask the second spectral data to form a third second spectral mask data. Furthermore, it should be noted that each parameter in the second mask matrix is either 0 or 1, allowing the second mask matrix and the second spectral data to be multiplied bitwise to obtain the corresponding second spectral mask data. Based on this, multiple second mask matrices can mask different components in the second spectral data, enabling separate attention to different components and thus improving the accuracy of semantic extraction.
[0066] Step S142c: Mining attention features for each of the second spectral mask data to form the second spectral attention feature corresponding to each of the second spectral mask data.
[0067] In this embodiment of the application, after forming the second spectral mask data, attention features can be mined for each second spectral mask data to form a second spectral attention feature corresponding to each second spectral mask data. For example, for the first second spectral mask data, convolution processing can be performed on the second spectral mask data to obtain the corresponding convolution feature. Then, self-attention processing can be performed on the convolution feature to obtain the second spectral attention feature corresponding to the first second spectral mask data.
[0068] Step S142d: Fuse multiple second spectral attention features to form target audio semantic features.
[0069] In this embodiment of the application, after forming the second spectral attention feature, multiple second spectral attention features can be fused to form a target audio semantic feature; for example, multiple second spectral attention features can be added or averaged to form a target audio semantic feature, or multiple second spectral attention features can be concatenated, and then the concatenated features can be sequentially processed by convolution, pooling and fully connected to form a target audio semantic feature.
[0070] In addition, since the data is first converted to the frequency domain in steps S141 and S142, and then the semantic extraction operation is performed, the resulting target image semantic features and target audio semantic features are both frequency domain features. Therefore, in subsequent steps, when performing semantic enhancement, the problem of reliability reduction caused by cross-modal feature processing can be avoided, so that the formed enhanced audio semantic features have high reliability and achieve effective representation of the target environment.
[0071] It is understood that in step S143 above, the specific method of performing semantic enhancement operation on the target audio semantic features is not limited and can be selected according to actual needs.
[0072] For example, in an alternative implementation, in order to fully integrate the semantic features of the target image into the semantic features of the target audio, thereby fully enhancing the semantic features of the target audio and capturing complex semantic relationships, the above step S143 may further include steps S143a, S143b, S143c and S143d, the specific contents of each step are as follows.
[0073] Step S143a: Based on the feature size of the semantic features of the target image and the feature size of the semantic features of the target audio, extract a first enhancement sub-matrix whose feature size matches the feature size of the semantic features of the target image from the target enhancement matrix, and extract a second enhancement sub-matrix whose feature size matches the feature size of the semantic features of the target audio.
[0074] In this embodiment, based on the feature size of the target image semantic features and the feature size of the target audio semantic features, a first enhanced sub-matrix whose feature size matches the feature size of the target image semantic features can be extracted from the target enhanced matrix, and a second enhanced sub-matrix whose feature size matches the feature size of the target audio semantic features can also be extracted. That is, the feature size of the first enhanced sub-matrix is the same as the feature size of the target image semantic features, and the feature size of the second enhanced sub-matrix is the same as the feature size of the target audio semantic features. Specifically, the target enhanced matrix can be decomposed to form a first enhanced sub-matrix and a second enhanced sub-matrix. The product of the first and second enhanced sub-matrixes is the target enhanced matrix. For example, if the feature size of the target enhanced matrix is n*n, it can be decomposed into two sub-matrices, n*1 and 1*n. The n*1 sub-matrix can be determined as the first enhanced sub-matrix, and the 1*n sub-matrix can be transposed to obtain the n*1 transpose matrix, which is then determined as the second enhanced sub-matrix. Based on this, the number of parameters stored and tuned can be reduced during storage and parameter tuning.
[0075] Step S143b: Based on the first enhancement sub-matrix, perform linear mapping on the semantic features of the target image to form target image mapping features.
[0076] In this embodiment, after extracting the first enhancement sub-matrix, the semantic features of the target image can be linearly mapped based on the first enhancement sub-matrix to form target image mapping features. For example, the first enhancement sub-matrix and the semantic features of the target image can be multiplied to achieve a linear mapping of the semantic features of the target image, that is, to capture linear semantic relationships, thereby obtaining target image mapping features.
[0077] Step S143c: Based on the second enhancement sub-matrix, perform linear mapping on the target audio semantic features to form target audio mapping features.
[0078] In this embodiment, after extracting the second enhancement sub-matrix, the target audio semantic features can be linearly mapped based on the second enhancement sub-matrix to form target audio mapping features. For example, the second enhancement sub-matrix can be multiplied with the target audio semantic features to achieve a linear mapping of the target audio semantic features, that is, to capture linear semantic relationships, thereby obtaining the target audio mapping features.
[0079] Step S143d: Based on the target image mapping features, perform cross-attention processing and residual connection on the target audio mapping features to form enhanced audio semantic features.
[0080] In this embodiment, after forming the target image mapping feature and the target audio mapping feature, cross-attention processing and residual connections can be performed on the target audio mapping feature based on the target image mapping feature to form enhanced audio semantic features. That is, cross-attention processing can first be performed on the target audio mapping feature based on the target image mapping feature to obtain corresponding cross-attention features, thus fusing the semantic features of the audio and image dimensions. Then, the cross-attention features and the target audio mapping feature can be connected to form enhanced audio semantic features. In one embodiment, the cross-attention features and the target audio mapping feature can be added or averaged to obtain the enhanced audio semantic features. In another embodiment, the cross-attention features and the target audio mapping feature can be concatenated, and then the concatenated features can be processed by convolution, pooling, and fully connected layers to obtain enhanced audio semantic features. In another implementation, the cross-attention features can be linearly mapped to obtain corresponding linearly mapped features. Then, non-linear activation (such as through the sigmoid function) can be applied to the linearly mapped features to obtain the corresponding linear activation parameter distribution. Then, the linear activation parameter distribution and the target audio mapping features can be multiplied bitwise to obtain enhanced audio semantic features. In this way, by using gating filtering, the cross-attention features and the target audio mapping features can be fully integrated, enabling the mining of more complex semantic relationships.
[0081] For example, in another alternative implementation, in order to enable multi-level fusion of target image semantic features and target audio semantic features through semantic enhancement operations, thereby improving the reliability of semantic enhancement operations and making the semantic representation ability of the resulting enhanced audio semantic features better, the specific contents of each of the above steps S143e, S143f, S143g, and S143h are as follows (in conjunction with...). Figure 5 (As shown).
[0082] Step S143e: Perform compression operations on the target image semantic features and the target audio semantic features respectively to form target image compression features and target audio compression features with multiple feature sizes.
[0083] In this embodiment, compression operations (i.e., downsampling) are performed on the target image semantic features and the target audio semantic features respectively, forming target image compressed features and target audio compressed features of multiple feature sizes. That is, multiple compression operations are performed on the target image semantic features to form target image compressed features of multiple feature sizes, wherein the target image compressed feature with the largest feature size is the target image semantic feature. Similarly, multiple compression operations are performed on the target audio semantic features to form target audio compressed features of multiple feature sizes, wherein the target audio compressed feature with the largest feature size is the target audio semantic feature.
[0084] In step S143f, in the first stage, based on the first enhancement matrix included in the target enhancement matrix and the target image compression feature with the largest feature size, a semantic enhancement operation is performed on the target audio compression feature with the largest feature size to form the enhancement feature of the first stage.
[0085] In this embodiment, after forming target image compression features and target audio compression features with multiple feature sizes, multiple stages of semantic enhancement operations can be performed based on the compression features with multiple feature sizes. Specifically, in the first stage, based on the first enhancement matrix included in the target enhancement matrix and the target image compression feature with the largest feature size, a semantic enhancement operation is performed on the target audio compression feature with the largest feature size to form the enhanced feature of the first stage. There is a one-to-one correspondence between the multiple feature sizes and the number of enhancement stages; for example, the first stage corresponds to the largest feature size, and the last stage corresponds to the smallest feature size.
[0086] In step S143g, in each of the second and subsequent stages, the enhancement features of the previous stage are connected to the target audio compression features of the corresponding feature size, and semantic enhancement is performed on the connection result based on the enhancement matrix of the current stage and the target image compression features of the corresponding feature size included in the target enhancement matrix, to form the enhancement features of the current stage.
[0087] In this embodiment, after completing the first stage of enhancement, in each subsequent stage, the enhanced features of the previous stage can be connected to the target audio compression features of the corresponding feature size. Furthermore, based on the enhancement matrix of the current stage included in the target enhancement matrix and the target image compression features of the corresponding feature size, a semantic enhancement operation is performed on the connection result to form the enhanced features of the current stage. For example, for the second stage, the enhanced features of the first stage can be connected to the target audio compression features with the second largest feature size. Furthermore, based on the enhancement matrix of the second stage included in the target enhancement matrix and the target image compression features with the second largest feature size, a semantic enhancement operation is performed on the connection result to form the enhanced features of the second stage. This process can be repeated to form the enhanced features of the third and fourth stages, and so on.
[0088] Step S143h: Based on the enhancement features of the last stage, determine the enhanced audio semantic features.
[0089] In this embodiment of the application, after the enhancement features of the last stage are formed, the enhanced audio semantic features can be determined based on the enhancement features of the last stage. For example, the enhancement features of the last stage can be directly determined as the enhanced audio semantic features.
[0090] It is understood that the specific manner in which the enhanced features of the current stage are formed in step S143g described above is not limited. For example, in an alternative implementation, step S143g may further include the following: First, in each of the second and subsequent stages, the enhancement features of the previous stage and the target audio compression features of the corresponding feature size can be concatenated, and the concatenated result can be pooled and compressed to form the connection result of the current stage. After pooling and compression, the connection result of the current stage has the same feature size as the enhancement matrix and the target image compression features of the corresponding feature size of the current stage. Secondly, based on the current stage enhancement matrix and the target image compression features of the corresponding feature size included in the target enhancement matrix, semantic enhancement operations can be performed on the connection results of the current stage (the specific enhancement process can be referred to the relevant description above, such as steps S143a-143d, i.e., from the two enhancement sub-matrices in the enhancement matrix), to form the enhancement features of the current stage.
[0091] Combination Figure 5 This application also provides a data analysis-based environmental atmosphere control device applicable to the aforementioned electronic devices. The data analysis-based environmental atmosphere control device may include a control scheme acquisition module, an audio data acquisition module, an image data acquisition module, a semantic enhancement module, and a control scheme determination module.
[0092] The control scheme acquisition module is used to acquire multiple candidate atmosphere control schemes pre-configured for a target song, wherein the target song is the song currently being sung by the target user in the target environment, and the candidate atmosphere control schemes at least include schemes for controlling the lighting in the target environment. In this embodiment, the control scheme acquisition module can be used to execute... Figure 2 The relevant content regarding the control scheme acquisition module in step S110 shown can be found in the previous description of step S110.
[0093] The audio data acquisition module is used to acquire target audio data in the target environment, wherein the target audio data includes the sound of the target user singing the target song and background sound. In this embodiment of the application, the audio data acquisition module can be used to perform... Figure 2 The relevant content regarding the audio data acquisition module in step S120 shown can be found in the previous description of step S120.
[0094] The image data acquisition module is used to acquire target image data in the target environment, wherein the target image data at least reflects the pose information of the target user. In this embodiment, the image data acquisition module can be used to perform... Figure 2 The relevant content regarding the image data acquisition module in step S130 shown can be found in the preceding description of step S130.
[0095] The semantic enhancement module is used to perform semantic enhancement operations on the semantic information of the target audio data based on the semantic information of the target image data, forming enhanced audio semantic features. In these enhanced audio semantic features, the semantic information of the target audio data is primary, and the semantic information of the target image data is secondary. In this embodiment, the semantic enhancement module can be used to execute... Figure 2 The relevant content regarding the semantic enhancement module in step S140 shown can be found in the previous description of step S140.
[0096] The control scheme determination module is used to determine a target atmosphere control scheme that matches the target environment from among the multiple candidate atmosphere control schemes based on the enhanced audio semantic features. The target atmosphere control scheme is used to control the lighting in the target environment. In this embodiment, the control scheme determination module can be used to execute... Figure 2 The relevant content regarding the control scheme determination module in step S150 shown can be found in the previous description of step S150.
[0097] In this embodiment of the application, corresponding to the above-described environmental atmosphere control method based on data analysis applied to the electronic device, a computer-readable storage medium is also provided, which stores a computer program that executes the various steps of the environmental atmosphere control method based on data analysis when the computer program is run.
[0098] The steps executed by the aforementioned computer program during runtime will not be described in detail here, but can be found in the explanation of the data analysis-based environmental atmosphere control method described above.
[0099] In summary, the environmental atmosphere control method, apparatus, device, and medium based on data analysis provided in this application first acquire multiple candidate atmosphere control schemes pre-configured for a target song; second, acquire target audio data in the target environment; then, acquire target image data in the target environment; subsequently, based on the semantic information of the target image data, perform semantic enhancement operations on the semantic information of the target audio data to form enhanced audio semantic features; finally, based on the enhanced audio semantic features, determine the target atmosphere control scheme that matches the target environment from among the multiple candidate atmosphere control schemes. Based on the above, on the one hand, since the target atmosphere control scheme is determined from multiple candidate atmosphere control schemes, the range of determination for the target atmosphere control scheme is relatively small, thus ensuring the accuracy of determining the target atmosphere control scheme. On the other hand, since semantic enhancement is performed on the target audio data based on the target image data containing the target user's posture information, the accuracy and richness of the representation of the environmental atmosphere by the enhanced audio semantic features are improved, thus ensuring the reliability of the determined target atmosphere control scheme and enabling reliable control of the environmental atmosphere, thereby improving the relatively poor environmental atmosphere control effect in existing technologies.
[0100] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0101] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0102] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. In the absence of further restrictions, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0103] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An environmental atmosphere control method based on data analysis, characterized in that, include: Obtain multiple candidate atmosphere control schemes pre-configured for a target song, wherein the target song is the song currently being sung by the target user in the target environment, and the candidate atmosphere control schemes include at least a scheme for controlling the lighting in the target environment; Acquire target audio data in the target environment, wherein the target audio data includes the sound of the target user singing the target song and background sound; Acquire target image data in the target environment, wherein the target image data is used to reflect at least the posture information of the target user; The target image data is frequency domain transformed to form first spectrum data. Multiple first mask matrices are used to mask the first spectrum data, forming multiple first spectrum mask matrices, which are learned from sample audio data, sample image data, and sample atmosphere control schemes. Attention features are mined for each first spectrum mask data, forming a first spectrum attention feature corresponding to each first spectrum mask data. Multiple first spectrum attention features are fused to form target image semantic features. Semantic extraction is performed on the target audio data to form target audio semantic features. Based on the target enhancement matrix and the target image semantic features, semantic enhancement is performed on the target audio semantic features to form enhanced audio semantic features. The target enhancement matrix is learned from sample audio data, sample image data, and sample atmosphere control schemes, and the initial target enhancement matrix is a random matrix. In the enhanced audio semantic features, the semantic information of the target audio data is primary, and the semantic information of the target image data is secondary. Based on the enhanced audio semantic features, a target atmosphere control scheme that matches the target environment is determined from among the multiple candidate atmosphere control schemes. The target atmosphere control scheme is used to control the lighting in the target environment.
2. The environmental atmosphere control method based on data analysis according to claim 1, characterized in that, The step of performing semantic enhancement operations on the target audio semantic features based on the target enhancement matrix and the target image semantic features to form enhanced audio semantic features includes: Based on the feature size of the semantic features of the target image and the feature size of the semantic features of the target audio, a first enhancement sub-matrix whose feature size matches the feature size of the semantic features of the target image is extracted from the target enhancement matrix, and a second enhancement sub-matrix whose feature size matches the feature size of the semantic features of the target audio is extracted. Based on the first enhancement sub-matrix, the semantic features of the target image are linearly mapped to form target image mapping features; Based on the second enhancement sub-matrix, the target audio semantic features are linearly mapped to form target audio mapping features; Based on the target image mapping features, cross-attention processing and residual connections are performed on the target audio mapping features to form enhanced audio semantic features.
3. The environmental atmosphere control method based on data analysis according to claim 1, characterized in that, The step of performing semantic enhancement operations on the target audio semantic features based on the target enhancement matrix and the target image semantic features to form enhanced audio semantic features includes: The semantic features of the target image and the semantic features of the target audio are compressed to form target image compressed features and target audio compressed features of multiple feature sizes; In the first stage, based on the first enhancement matrix included in the target enhancement matrix and the target image compression feature with the largest feature size, a semantic enhancement operation is performed on the target audio compression feature with the largest feature size to form the enhancement feature of the first stage. There is a one-to-one correspondence between the multiple feature sizes and the number of enhancement stages. In each of the second and subsequent stages, the enhancement features of the previous stage are connected to the target audio compression features of the corresponding feature size, and semantic enhancement is performed on the connection result based on the enhancement matrix of the current stage and the target image compression features of the corresponding feature size included in the target enhancement matrix, to form the enhancement features of the current stage. Based on the enhancement features of the last stage, the enhanced audio semantic features are determined.
4. The environmental atmosphere control method based on data analysis according to claim 3, characterized in that, The steps of connecting the enhancement features of the previous stage to the target audio compression features of the corresponding feature size in each of the second and subsequent stages, and performing semantic enhancement operations on the connection results based on the enhancement matrix of the current stage and the target image compression features of the corresponding feature size included in the target enhancement matrix, to form the enhancement features of the current stage, include: In each of the second and subsequent stages, the enhancement features from the previous stage and the target audio compression features of the corresponding feature size are concatenated, and the concatenation result is pooled and compressed to form the connection result of the current stage. Based on the target enhancement matrix, which includes the enhancement matrix of the current stage and the target image compression features of the corresponding feature size, semantic enhancement operations are performed on the connection results of the current stage to form the enhancement features of the current stage.
5. The environmental atmosphere control method based on data analysis according to any one of claims 1-4, characterized in that, The step of performing semantic extraction on the target audio data to form target audio semantic features includes: The target audio data is frequency domain converted to form second spectrum data; Multiple second mask matrices are used to mask the second spectral data to form multiple second spectral mask data. The multiple second mask matrices are formed by learning from sample audio data, sample image data and sample atmosphere control scheme. Attention features are mined for each of the second spectral mask data to form the second spectral attention features corresponding to each of the second spectral mask data; Multiple second-spectrum attention features are fused to form target audio semantic features.
6. An environmental atmosphere control device based on data analysis, characterized in that, include: The control scheme acquisition module is used to acquire multiple candidate atmosphere control schemes pre-configured for a target song, wherein the target song is the song currently being sung by the target user in the target environment, and the candidate atmosphere control schemes include at least a scheme for controlling the lights in the target environment. An audio data acquisition module is used to acquire target audio data in the target environment, wherein the target audio data includes the sound of the target user singing the target song and background sound; An image data acquisition module is used to acquire target image data in the target environment, wherein the target image data is used to reflect at least the posture information of the target user; A semantic enhancement module is used to perform frequency domain transformation on the target image data to form first spectrum data; to perform masking processing on the first spectrum data using multiple first mask matrices to form multiple first spectrum mask data, wherein the multiple first mask matrices are learned from sample audio data, sample image data, and sample atmosphere control schemes; to mine attention features for each first spectrum mask data to form a first spectrum attention feature corresponding to each first spectrum mask data; to fuse multiple first spectrum attention features to form target image semantic features; to perform semantic extraction on the target audio data to form target audio semantic features; and to perform semantic enhancement on the target audio semantic features based on the target enhancement matrix and the target image semantic features to form enhanced audio semantic features, wherein the target enhancement matrix is learned from sample audio data, sample image data, and sample atmosphere control schemes, and the initial target enhancement matrix is a random matrix; wherein the enhanced audio semantic features are based primarily on the semantic information of the target audio data and secondarily on the semantic information of the target image data. The control scheme determination module is used to determine, based on the enhanced audio semantic features, a target atmosphere control scheme that matches the target environment from among the multiple candidate atmosphere control schemes, wherein the target atmosphere control scheme is used to control the lighting in the target environment.
7. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor connected to the memory is used to execute the computer program stored in the memory to implement the environmental atmosphere control method based on data analysis as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains a computer program that, when executed, performs the environmental atmosphere control method based on data analysis as described in any one of claims 1-5.
Citation Information
Patent Citations
Automatic identification laser special effect system of Karaoke
CN103977576A
Audio noise reduction method and device, electronic equipment and storage medium
CN119964540A