Scene effect information acquisition method and device, equipment and storage medium

By semantic analysis of the text of audiobook audio, determining the scene type and obtaining matching scene effect information, the problem of how to improve the sense of audiobook image is solved and a richer audio experience is achieved.

CN119990133APending Publication Date: 2025-05-13GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311452400.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

How to improve the sense of picture of audio books to meet readers' needs for improving audio experience.

Method used

By performing semantic analysis of the text corresponding to the first audio, the scene type of the audio is determined, and the target scene effect information matching the scene type is obtained, including the target scene sound effects and/or the visual target multimedia content.

Benefits of technology

It enhances the sense of picture of the audio and improves the experience quality of audio content such as audio books, so that readers can more fully feel the scenes and atmosphere in the audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990133A_ABST
    Figure CN119990133A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a scene effect information acquisition method and device, equipment and a storage medium, and the method can comprise the steps: carrying out the semantic analysis of a first text corresponding to a first audio, and determining a scene type corresponding to the first audio; and obtaining target scene effect information matched with the scene type, wherein the target scene effect information comprises a target scene sound effect and / or visual target multimedia content. By implementing the method, the picture feeling of the audio can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of electronic equipment, and in particular to a method, device, equipment and storage medium for acquiring scene effect information. Background Art

[0002] In recent years, with the development of the Internet and changes in people's reading habits, audiobooks have received more and more attention from the market and are being accepted by more and more readers. How to improve the visual quality of audiobooks has become an urgent problem that the industry needs to solve. Summary of the invention

[0003] The present invention provides a method, device, equipment and storage medium for obtaining scene effect information, which can enhance the picture sense of audio. 。

[0004] A first aspect of an embodiment of the present application provides a method for acquiring scene effect information, comprising:

[0005] Performing semantic analysis on a first text corresponding to a first audio to determine a scene type corresponding to the first audio;

[0006] Target scene effect information matching the scene type is acquired, where the target scene effect information includes target scene sound effects and / or visualized target multimedia content.

[0007] A second aspect of an embodiment of the present application provides a device for acquiring scene effect information, including:

[0008] A scene type identification unit, configured to perform semantic analysis on a first text corresponding to a first audio to determine a scene type corresponding to the first audio;

[0009] The effect information acquisition unit is used to acquire target scene effect information matching the scene type, wherein the target scene effect information includes target scene sound effects and / or visualized target multimedia content.

[0010] A third aspect of the embodiments of the present application provides an electronic device, including:

[0011] A memory storing executable program code;

[0012] and a processor coupled to the memory;

[0013] The processor calls the executable program code stored in the memory, and when the executable program code is executed by the processor, the processor implements the method described in the first aspect of the embodiment of the present application.

[0014] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium having executable program code stored thereon. When the executable program code is executed by a processor, the method described in the first aspect of the embodiment of the present application is implemented.

[0015] A fifth aspect of an embodiment of the present application discloses a computer program product. When the computer program product runs on a computer, the computer executes the method described in the first aspect of the embodiment of the present application.

[0016] A sixth aspect of an embodiment of the present application discloses an application publishing platform, which is used to publish a computer program product. When the computer program product runs on a computer, the computer executes the method described in the first aspect of the embodiment of the present application.

[0017] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0018] In an embodiment of the present application, a semantic analysis is performed on a first text corresponding to a first audio to determine a scene type corresponding to the first audio; and target scene effect information matching the scene type is obtained, the target scene effect information including target scene sound effects and / or visualized target multimedia content.

[0019] By implementing this method, semantic analysis can be performed on the first text corresponding to the first audio to determine the scene type corresponding to the first audio, and target scene effect information matching the scene type can be obtained, which is conducive to enhancing the visual sense of the audio.

[0020] Among them, when the first audio is the audio corresponding to the audio book, this method can achieve the purpose of enhancing the visual sense of the audio book. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments and the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can be obtained based on these drawings.

[0022] Figure 1 It is a scene diagram of the method for obtaining scene effect information disclosed in the embodiment of the present application;

[0023] Figure 2A It is a flowchart of the method for obtaining scene effect information disclosed in the embodiment of the present application;

[0024] Figure 2B It is an interface diagram of the audiobook application disclosed in the embodiment of the present application;

[0025] Figure 3is another flowchart of the method for obtaining scene effect information disclosed in the embodiment of the present application;

[0026] Figure 4 It is another flowchart of the method for obtaining scene effect information disclosed in the embodiment of the present application;

[0027] Figure 5 It is another flowchart of the method for obtaining scene effect information disclosed in the embodiment of the present application;

[0028] Figure 6 It is another flowchart of the method for obtaining scene effect information disclosed in the embodiment of the present application;

[0029] Figure 7 It is a structural diagram of a device for acquiring scene effect information disclosed in an embodiment of the present application;

[0030] Figure 8 It is a structural diagram of an electronic device disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0031] The embodiments of the present application provide a method, device, equipment and storage medium for obtaining scene effect information, which can enhance the picture sense of audio.

[0032] In order to make the technical personnel in the technical field better understand the scheme of the present application, the technical scheme in the embodiment of the present application will be described below in conjunction with the drawings in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, not all of the embodiments. Based on the embodiments in the present application, they should all fall within the scope of protection of the present application.

[0033] It should be noted that, in this application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplarily" or "for example" in this application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplarily" or "for example" is intended to present related concepts in a specific way.

[0034] "At least one" means one or more, and "plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can mean: a, or b, or c, or a and b, or a and c, or b and c, or a, b and c, where a, b, c can be single or plural.

[0035] In recent years, with the development of the Internet and changes in people's reading habits, audiobooks have received more and more attention from the market and are being accepted by more and more readers. How to improve the visual quality of audiobooks has become an urgent problem that the industry needs to solve.

[0036] See also Figure 1 , Figure 1 It is a scene diagram of the method for obtaining scene effect information disclosed in the embodiment of this application. Figure 1 The illustrated scene diagram includes an electronic device 10. The electronic device 10 is installed with an audiobook application, which is used to play audiobooks (such as electronic books, crosstalk or stories, etc.), and a user can listen to audiobooks through the audiobook application.

[0037] In the technical solution of the present application, the electronic device 10 can perform semantic analysis on the first text corresponding to the first audio, determine the scene type corresponding to the first audio, and obtain target scene effect information matching the scene type to enhance the visual sense of the first audio. It can be understood that when the first audio is the audio corresponding to an audiobook, this method can achieve the purpose of enhancing the visual sense of the audiobook.

[0038] Among them, the electronic device 10 may include general handheld screen electronic devices, such as mobile phones, smart phones, portable terminals, terminals, personal digital assistants (PDA), portable multimedia players (Personal Media Player, PMP) devices, laptop computers, notebooks (Note Pad), wireless broadband (Wireless Broadband, Wibro) terminals, tablet computers (Personal Computer, PC), smart PCs, sales terminals (Point of Sales, POS) and car computers, etc.

[0039] The electronic device 10 may also include a wearable device. A wearable device is a portable electronic device that can be worn directly on the user or integrated into the user's clothes or accessories. Wearable devices are not only hardware devices, but can also realize powerful intelligent functions through software support, data interaction, and cloud server interaction, such as: computing function, positioning function, and alarm function. At the same time, they can also connect to mobile phones and various terminals. Wearable devices may include but are not limited to watches supported by wrists (such as watches, wrists, etc.), shoes supported by feet (such as shoes, socks or other products worn on the legs), glasses supported by the head (such as glasses, helmets, headbands, etc.), and various non-mainstream product forms such as smart clothing, school bags, crutches, accessories, etc.

[0040] The present application scheme is further described below in conjunction with specific embodiments:

[0041] See also Figure 2A , Figure 2A FIG. 1 is a flowchart of a method for obtaining scene effect information disclosed in an embodiment of the present application. Figure 2A The method for obtaining scene effect information shown may include the following steps:

[0042] 201. An electronic device performs semantic analysis on a first text corresponding to a first audio to determine a scene type corresponding to the first audio.

[0043] Among them, the first audio may be, in addition to the audio corresponding to the audiobook, the real-time audio of the host on the live broadcast platform, or the audio corresponding to the video screen of the video playback application, etc., which is not limited in the embodiments of the present application.

[0044] In an embodiment of the present application, the first text corresponding to the first audio may be pre-stored on the electronic device, or may be obtained by the electronic device by converting the first audio into text, which is not limited in the embodiment of the present application.

[0045] Exemplarily, if the first audio is the audio corresponding to an audiobook, the electronic device can pre-store the text of the audiobook, that is, the first text of the first audio; if the first audio is the real-time audio of the host, the electronic device only stores the audio, and the first text corresponding to the first audio requires the electronic device to convert the first audio into text.

[0046] The scene type may refer to the atmosphere type corresponding to the first audio. For example, the scene type may include a horror type, a joke type, a happy type, and a sad type.

[0047] In some embodiments, the electronic device performs semantic analysis on the first text corresponding to the first audio, which may mean that the electronic device determines whether the first text corresponding to the first audio contains a keyword, and if the keyword is contained, determines the scene type corresponding to the first audio according to the keyword. The electronic device may pre-store multiple scene types and keywords matching each scene type.

[0048] Exemplarily, the keywords include "happy", "sad", and "hello", wherein the scene types corresponding to "happy" and "hello" are happy scenes, and the scene type corresponding to "sad" is a sad scene.

[0049] In some embodiments, the electronic device performs semantic analysis on the first text corresponding to the first audio to determine the scene type corresponding to the first audio, which may include:

[0050] Added operation for electronic equipment detection effect;

[0051] In response to the above-mentioned effect adding operation, the electronic device performs semantic analysis on the first text corresponding to the first audio and determines the scene type corresponding to the first audio.

[0052] The effect adding operation may include but is not limited to at least one of the following: voice operation, gesture operation, and touch operation.

[0053] Exemplarily, when the effect adding operation includes a voice operation, the effect adding operation may be a voice input of “start effect adding”.

[0054] When the effect adding operation includes a gesture operation, the effect adding operation may be shaking the electronic device left or right.

[0055] When the effect adding operation includes a touch operation, the effect adding operation may be a touch control for triggering a virtual button for starting the effect adding, and the virtual button is on the interface of a corresponding application (for example, an audiobook application, a live broadcast application, or a video playback application). Figure 2B , Figure 2B This is an interface diagram of the audiobook application disclosed in the embodiment of this application. Figure 2B The interface diagram shown includes a virtual button 210 for triggering the addition of a start effect.

[0056] It can be understood that the effect adding operation is used to trigger the electronic device to start the scene effect adding function of the corresponding application.

[0057] In some embodiments, the electronic device can also determine whether the scene effect adding function of the specified application is turned on when detecting that the specified application is started. If the scene effect adding function is not turned on, it outputs instruction information for instructing the user to turn on the scene effect adding function of the specified application (that is, input the above-mentioned effect adding operation).

[0058] In the embodiment of the present application, the above indication information may include but is not limited to at least one of the following: text, voice, animation, etc.

[0059] The designated application may be an application that has a scene effect adding function. For example, the designated application may include an audiobook application, a live broadcast application, and a video playback application.

[0060] By implementing this method, the electronic device can also detect whether the scene effect adding function is turned on or not when a specified application is started, and if the scene effect adding function is not turned on, output instruction information for instructing the user to turn on the scene effect adding function, so as to remind the user to turn on the scene effect adding function in time.

[0061] In some embodiments, if the scene effect adding function is not turned on, the electronic device may further output guidance information for turning on the scene effect adding function to guide the user to quickly turn on the scene effect adding function.

[0062] 202. The electronic device obtains target scene effect information matching the scene type, where the target scene effect information includes target scene sound effects and / or visualized target multimedia content.

[0063] Among them, scene sound effects refer to the effects created by sound, which are noises or sounds added to the vocal cords to enhance the realism, atmosphere or dramatic message of the scene. Visual multimedia content can include but is not limited to animations and emoticons.

[0064] In an embodiment of the present application, the electronic device may pre-store a scene effect information library, which may include multiple scene types and scene effect information corresponding to each scene type. The electronic device obtains target scene effect information matching the scene type, which may include: the electronic device searches the scene effect information library for target scene effect information matching the scene type corresponding to the first audio.

[0065] In some embodiments, the scene effect information library can support updating, and the electronic device can request the latest scene effect information library from the cloud server in response to the effect information update operation. The effect information update operation may include but is not limited to at least one of the following: voice operation, gesture operation, and touch operation.

[0066] By implementing the above method, the electronic device can determine the scene type corresponding to the first audio by performing semantic analysis on the first text corresponding to the first audio, and obtain target scene effect information matching the scene type, thereby facilitating enhancing the visual sense of the audio (audiobooks, real-time audio of the host, and audio corresponding to the video screen).

[0067] See also Figure 3 , Figure 3 FIG. 1 is another flowchart of the method for obtaining scene effect information disclosed in the embodiment of the present application. Figure 3 The method for obtaining scene effect information shown may include the following steps:

[0068] 301. The electronic device converts the first audio into text to obtain a first text corresponding to the first audio.

[0069] In some embodiments, the electronic device converts the first audio into text to obtain a first text corresponding to the first audio, which may include:

[0070] The electronic device extracts a sound feature of the first audio; wherein the sound feature may include but is not limited to information such as frequency, pitch, volume, etc.;

[0071] The electronic device constructs a language model and an acoustic model according to the sound features; the acoustic model may refer to the mapping of speech features to phonemes, and the language model may refer to the mapping between words and words, or between words and sentences;

[0072] The electronic device obtains a first text corresponding to the first audio according to the language model and the acoustic model.

[0073] 302. The electronic device performs semantic analysis on the first text to determine a scene type corresponding to the first audio.

[0074] In some embodiments, the electronic device performs semantic analysis on the first text to determine the scene type corresponding to the first audio, which may include:

[0075] The electronic device obtains a second text corresponding to a second audio before the first audio;

[0076] The electronic device performs semantic analysis on the first text based on the second text to determine the scene type corresponding to the first audio.

[0077] Among them, the second audio before the first audio refers to the audio that occurs before the time when the first audio occurs. Exemplarily, when the first audio is an audio segment of an audio book to be output, the second audio is one or more audio segments of the audio book that were previously output. When the first audio is a real-time audio segment of the host to be output, the second audio is one or more real-time audio segments that were previously output. When the first audio is the audio corresponding to the video screen to be output, the second audio is the audio corresponding to one or more video screens that were previously output.

[0078] It is understandable that the electronic device can greatly improve the accuracy of scene recognition by performing semantic analysis on the first text in combination with the above.

[0079] In some embodiments, the electronic device performs semantic analysis on the first text corresponding to the first audio based on the second text to determine the scene type corresponding to the first audio, which may include: the electronic device performs text classification, sentence similarity processing and entity recognition on the first text and the second text respectively to determine the scene type corresponding to the first audio.

[0080] Among them, the electronic device can use the chat robot model to perform text classification, sentence similarity processing and entity recognition on the first text and the second text respectively, and determine the scene type corresponding to the first audio.

[0081] Exemplarily, the chatbot model is ChatGPT (Chat Generative Pre-trained Transformer).

[0082] 303. The electronic device obtains a target scene sound effect that matches the scene type.

[0083] The scene effect information library mentioned in step 202 may include scene sound effects corresponding to each scene type. There may be one or more scene sound effects corresponding to each scene type, which is not limited in the embodiment of the present application.

[0084] When there is one scene sound effect corresponding to each scene type, the scene effect information library may include a scene sound effect library, and the scene sound effect library may include the scene sound effect corresponding to each scene type.

[0085] When each scene type corresponds to multiple scene sound effects, the scene effect information library may include a scene sound effect library corresponding to each scene type, and the scene sound effect library corresponding to each scene type includes multiple scene sound effects corresponding to each scene type. The electronic device obtains the target scene sound effect matching the scene type, which may include:

[0086] The electronic device obtains a scene sound effect library matching the scene type;

[0087] The electronic device obtains the target scene sound effect from a scene sound effect library that matches the scene type.

[0088] In some embodiments, the electronic device obtains the target scene sound effect from the scene sound effect library matching the scene type, which may include:

[0089] Electronic devices obtain user information;

[0090] The electronic device obtains the target scene sound effect matching the user information from the scene sound effect library matching the scene type.

[0091] The user information may include but is not limited to at least one of the following: gender, age, hobbies, occupation, etc. By implementing the method, the electronic device can obtain scene sound effects that match the user information, thereby realizing the addition of personalized scene sound effects.

[0092] In some embodiments, the electronic device obtains the target scene sound effect from the scene sound effect library matching the scene type, which may include:

[0093] The electronic device uses the locked scene sound effect in the scene sound effect library that matches the scene type as the target scene sound effect.

[0094] The locked scene sound effect refers to the scene sound effect manually selected by the user in advance. It can be understood that the electronic device can support manual selection of scene sound effects of various scene types.

[0095] Exemplarily, when the first audio is the real-time audio of the host on the live broadcast platform, the electronic device performs semantic analysis on the real-time audio in combination with the context and determines that the real-time audio is the beginning of a joke. The scene type at this time is a joke type, and the electronic device selects the target scene sound effect from the scene sound effect library that matches the joke type.

[0096] Exemplarily, when the first audio is the real-time audio of the anchor on the live broadcast platform, the electronic device performs semantic analysis on the real-time audio in combination with the context and determines that the real-time audio is the beginning of a horror ghost story. The scene type at this time is the horror type, and the electronic device selects the target scene sound effect from the scene sound effect library that matches the horror type.

[0097] 304. The electronic device mixes the first audio with the target scene sound effect to obtain the target audio.

[0098] 305. The electronic device outputs the target audio.

[0099] In an embodiment of the present application, the electronic device includes an audio playback track and an audio mixer.

[0100] The electronic device mixes the first audio with the target scene sound effect to obtain the target audio, which may include: the electronic device mixes the first audio with the target scene sound effect through an audio mixer to obtain the target audio.

[0101] The electronic device outputting the target audio may include: the electronic device outputting the target audio through an audio playback track.

[0102] It should be noted that the above audio track is newly created by the electronic device and is different from the original audio track.

[0103] By implementing the above method, the electronic device can determine the scene type corresponding to the first audio by performing semantic analysis on the first text corresponding to the first audio, and obtain the target scene sound effect matching the scene type, as well as mix the target scene sound effect and the first audio to obtain the target audio, and output the target audio, thereby achieving the effect of adding rich ambient sound to the original monotonous audio, optimizing the audience experience and increasing the fun.

[0104] Exemplarily, the first audio is the real-time audio of the host. Figure 4 , the process of adding scene sound effects is explained. Figure 4 The process of adding scene sound effects to the electronic device shown is as follows:

[0105] 401. The electronic device obtains the real-time audio of the host.

[0106] 402. The electronic device converts the real-time audio into text.

[0107] 403. The electronic device performs semantic analysis on the text corresponding to the real-time audio.

[0108] 404. The electronic device identifies the scene type based on the analysis result.

[0109] 405. The electronic device obtains a target scene sound effect matching the above scene type from a scene sound effect library.

[0110] 406. The electronic device mixes the target scene sound effect and the real-time audio to obtain the target audio.

[0111] 407. The electronic device outputs the target audio.

[0112] See also Figure 5 , Figure 5 FIG. 1 is another flowchart of the method for obtaining scene effect information disclosed in the embodiment of the present application. Figure 5 The method for obtaining scene effect information shown may include the following steps:

[0113] 501. The electronic device performs text conversion on a first audio to obtain a first text corresponding to the first audio.

[0114] 502. The electronic device performs semantic analysis on the first text to determine a scene type corresponding to the first audio.

[0115] For a detailed description of step 501 - step 502 , please refer to the introduction of step 301 - step 302 above, which will not be repeated here.

[0116] 503. The electronic device obtains visualized target multimedia content matching the scene type.

[0117] The scene effect information library mentioned in step 202 may include visual multimedia content corresponding to each scene type. The visual multimedia content corresponding to each scene type may be one or more, which is not limited in the embodiment of the present application.

[0118] When there is one visualized multimedia content corresponding to each scene type, the scene effect information library may include a visualized multimedia content library, and the visualized multimedia content library may include a visualized multimedia content library corresponding to each scene type.

[0119] When each scene type corresponds to multiple visual multimedia contents, the scene effect information library may include a visual multimedia content library corresponding to each scene type, and the visual multimedia content library corresponding to each scene type includes multiple visual multimedia contents corresponding to each scene type. The electronic device obtains target visual multimedia contents matching the scene type, which may include:

[0120] The electronic device obtains a visual multimedia content library matching the scene type;

[0121] The electronic device obtains target visualized multimedia content from a visualized multimedia content library matching the scene type.

[0122] In some embodiments, the electronic device obtains target visualized multimedia content from a visualized multimedia content library matching the scene type, which may include:

[0123] Electronic devices obtain user information;

[0124] The electronic device obtains target visualized multimedia content matching the user information from a visualized multimedia content library matching the scene type.

[0125] The user information may include but is not limited to at least one of the following: gender, age, hobbies, occupation, etc. By implementing the method, the electronic device can obtain visualized multimedia content matching the user information, thereby realizing the addition of personalized visualized multimedia content.

[0126] In some embodiments, the electronic device obtains target visualized multimedia content from a visualized multimedia content library matching the scene type, which may include:

[0127] The electronic device uses the locked visualized multimedia content in the visualized multimedia content library matching the scene type as the target visualized multimedia content.

[0128] The locked visualized multimedia content refers to the visualized multimedia content manually selected by the user in advance. It can be understood that the electronic device supports manual selection of visualized multimedia content of various scene types.

[0129] Exemplarily, when the first audio is the real-time audio of the host on the live broadcast platform, the electronic device performs semantic analysis on the real-time audio in combination with the context and determines that the real-time audio is the beginning of a joke. The scene type at this time is a joke type, and the electronic device selects visualized target multimedia content from a visualized multimedia content library that matches the joke type.

[0130] Exemplarily, when the first audio is the real-time audio of the host on the live broadcast platform, the electronic device performs semantic analysis on the real-time audio in combination with the context and determines that the real-time audio is the beginning of a horror ghost story. The scene type at this time is horror type, and the electronic device selects visual target multimedia content from the visual multimedia content library that matches the horror type.

[0131] 504. The electronic device outputs visualized target multimedia content.

[0132] In some embodiments, the electronic device outputs the visualized target multimedia content in the following ways, including but not limited to:

[0133] The electronic device may display the visualized target multimedia content on a display screen of the electronic device;

[0134] or,

[0135] The electronic device can project the visualized target multimedia content onto any plane.

[0136] It should be noted that the method of outputting visualized target multimedia content by projection by electronic devices can overcome the disadvantage that the user's viewing experience is affected by the small display screen of the electronic device (such as wearable devices).

[0137] In some embodiments, if the first audio is the audio of an audiobook, the visualized target multimedia content is displayed in full screen on the display screen of the electronic device. If the first audio is the real-time audio of the anchor or the corresponding audio of the video screen, the visualized target multimedia content is displayed on a partial area of ​​the display screen of the electronic device.

[0138] In some embodiments, the electronic device displays the visualized target multimedia content on a display screen of the electronic device, which may include:

[0139] When the first audio is the real-time audio of the host or the corresponding audio of the video screen, a split-screen operation is performed on the display screen of the electronic device to obtain a first screen and a second screen;

[0140] The electronic device displays the live broadcast picture / video picture on the first screen;

[0141] The electronic device displays the visualized target multimedia content on the second screen.

[0142] By implementing this method, when the first audio is the real-time audio of the host or the corresponding audio of the video screen, the visualized target multimedia content is displayed in a split screen with the live screen / video screen, which can avoid blocking the live screen / video screen and affecting the user's viewing experience.

[0143] By implementing the above method, the electronic device can perform semantic analysis on the first text corresponding to the first audio, determine the scene type corresponding to the first audio, obtain visualized target multimedia content matching the scene type, and output the visualized target multimedia content, thereby achieving the purpose of enhancing the picture sense of the audio.

[0144] See also Figure 6 , Figure 6 FIG. 1 is another flowchart of the method for obtaining scene effect information disclosed in the embodiment of the present application. Figure 6 The method for obtaining scene effect information shown may include the following steps:

[0145] 601. The electronic device converts the first audio into text to obtain a first text corresponding to the first audio.

[0146] 602. The electronic device performs semantic analysis on the first text to determine a scene type corresponding to the first audio.

[0147] For a detailed explanation of step 601 - step 602 , please refer to the above description of step 301 - step 302 , which will not be repeated here.

[0148] 603. The electronic device obtains target scene sound effects and visualized target multimedia content that match the scene type.

[0149] For a detailed explanation of step 603, please refer to the above description of step 303 and step 503, which will not be repeated here.

[0150] 604. The electronic device mixes the first audio with the target scene sound effect to obtain the target audio.

[0151] For a detailed explanation of step 604, please refer to the above description of step 304, which will not be repeated here.

[0152] 605. The electronic device outputs the target audio and outputs the visualized target multimedia content.

[0153] It should be noted that the electronic device can output visual target multimedia content while outputting the target audio.

[0154] By implementing the above method, the electronic device can determine the scene type corresponding to the first audio by performing semantic analysis on the first text corresponding to the first audio, and obtain the target scene sound effects and visualized target multimedia content that match the scene type, as well as mix the target scene sound effects and the first audio to obtain the target audio, and output the target audio and visualized target multimedia content. In addition to achieving the effect of adding rich ambient sounds to the original monotonous audio, it can also achieve the effect of adding visualized ambient multimedia content, further enhancing the picture sense of the audio.

[0155] See also Figure 7 , Figure 7 FIG. 1 is a structural diagram of a device for acquiring scene effect information disclosed in an embodiment of the present application. Figure 7 The device shown may include a scene type identification unit 701 and an effect information acquisition unit 702; wherein:

[0156] A scene type identification unit 701 is used to perform semantic analysis on a first text corresponding to a first audio to determine a scene type corresponding to the first audio;

[0157] The effect information acquisition unit 702 is used to acquire target scene effect information matching the scene type, where the target scene effect information includes target scene sound effects and / or visualized target multimedia content.

[0158] In some embodiments, the scene type identification unit 701 is used to perform semantic analysis on the first text corresponding to the first audio, and the method of determining the scene type corresponding to the first audio may specifically include:

[0159] The scene type identification unit 701 is used to obtain a second text corresponding to a second audio before the first audio; and, based on the second text, perform semantic analysis on the first text corresponding to the first audio to determine the scene type corresponding to the first audio.

[0160] In some embodiments, Figure 7 The device shown may also include a speech-to-text unit ( Figure 7 (not shown), a speech-to-text unit is used to perform text conversion on the first audio to obtain a first text corresponding to the first audio before the scene type recognition unit 701 performs semantic analysis on the first text corresponding to the first audio to determine the scene type corresponding to the first audio.

[0161] In some embodiments, the target scene effect information includes target scene sound effects, Figure 7 The device shown may also include a sound mixing unit ( Figure 7(not shown), a sound effect mixing unit, which is used to mix the first audio and the target scene sound effect to obtain the target audio after the effect information acquisition unit 702 acquires the target scene effect information matching the scene type; and outputs the target audio.

[0162] In some embodiments, the target scene effect information includes the target scene sound effect, and the method used by the effect information acquisition unit 702 to acquire the target scene sound effect matching the scene type may specifically include:

[0163] The effect information acquisition unit 702 is used to acquire a scene sound effect library that matches the scene type; and to acquire a target scene sound effect from the scene sound effect library.

[0164] In some embodiments, the effect information acquisition unit 702 may be used to acquire the target scene sound effect from the scene sound effect library in the following manner:

[0165] The effect information acquisition unit 702 is used to acquire user information; and to acquire target scene sound effects matching the user information from a scene sound effect library.

[0166] In some embodiments, the target scene effect information includes visualized target multimedia content. The effect information acquisition unit 702 is further configured to output the visualized target multimedia content after acquiring the target scene effect information matching the scene type.

[0167] In some embodiments, the scene type identification unit 701 is used to perform semantic analysis on the first text corresponding to the first audio, and the method of determining the scene type corresponding to the first audio may specifically include:

[0168] The scene type identification unit 701 is used to detect an effect adding operation; and, in response to the effect adding operation, perform semantic analysis on a first text corresponding to a first audio to determine a scene type corresponding to the first audio.

[0169] See also Figure 8 , Figure 8 FIG. 1 is a structural diagram of an electronic device disclosed in an embodiment of the present application. Figure 8 The electronic device shown includes: a processor 810, a memory 820, a display unit 830, an input unit 840, a sensor 850, an audio circuit 860 and other components.

[0170] The processor 810 is the control center of the electronic device, which uses various interfaces and lines to connect various parts of the entire electronic device, and executes various functions of the electronic device and processes data by running or executing software programs and / or modules stored in the memory 820, and calling data stored in the memory 820, so as to monitor the electronic device as a whole. Optionally, the processor 810 may include one or more processing units; optionally, the processor 810 may integrate an application processor, which mainly processes operating devices, user interfaces, and application programs, etc. Of course, other processors may also be included, which are not listed here one by one.

[0171] The memory 820 can be used to store software programs and modules. The processor 810 executes various functional applications and data processing of the electronic device by running the software programs and modules stored in the memory 820. The memory 820 may mainly include a program storage area and a data storage area, wherein the program storage area may store operating devices, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device (such as audio data, phone book, etc.), etc. In addition, the memory 820 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0172] The display unit 830 can be used to display information input by the user or information provided to the user and various menus of the electronic device. The display unit 830 may include a display panel. Optionally, the display panel may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel may cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 810 to determine the type of touch event. Then, the processor 810 provides a corresponding visual output on the display panel according to the type of touch event. Among them, the touch panel and the display panel are not in contact with each other. Figure 8 The touch panel and the display panel can be used as two independent components to realize the input and output functions of the electronic device, or the touch panel and the display panel can be integrated to realize the input and output functions of the electronic device.

[0173] The input unit 840 can be used to receive input digital or character information, and to generate key signal input related to user settings and function control of the electronic device. Specifically, the input unit 840 may include a touch panel and other input devices. A touch panel, also known as a touch screen, can collect user touch operations on or near it (such as operations performed by a user using any suitable object or accessory such as a finger, stylus, etc. on or near the touch panel), and drive the corresponding connection device according to a pre-set program. In addition, a variety of types such as resistive, capacitive, infrared, and surface acoustic waves can be used to implement the touch panel. In addition to the touch panel, the input unit 840 may also include other input devices. Specifically, other input devices may include, but are not limited to, one or more of function keys (such as volume control keys, switch keys, etc.), trackballs, joysticks, etc.

[0174] The electronic device may also include at least one sensor 850, such as a magnetometer, a gyroscope sensor, a motion sensor, and other sensors. Specifically, the magnetometer is used to determine the orientation of the electronic device, and the gyroscope sensor can be used to determine the motion posture of the electronic device. It can be used for anti-shake shooting, and can also be used for navigation and somatosensory game scenes. As a type of motion sensor, the acceleration sensor can detect the magnitude of acceleration in all directions, and can detect the magnitude and direction of gravity when stationary. It can be used for applications that identify the posture of the electronic device, such as horizontal and vertical screen switching, related games, magnetometer posture calibration, etc.; as for other sensors that the electronic device can also be configured with, such as pressure gauges, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be repeated here.

[0175] The audio circuit 860 may include a speaker and a microphone, and may provide an audio interface between a user and an electronic device. The audio circuit 860 may transmit the electrical signal converted from the received audio data to the speaker, which is converted into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 860 and converted into audio data, and then the audio data is output to the processor 810 for processing, and then sent to another device through the video circuit, or the audio data is output to the memory 820 for further processing.

[0176] Although not shown, the electronic device may further include a power supply and a camera. Optionally, the camera may be located at the front or rear of the electronic device, which is not limited in the present embodiment.

[0177] It is to be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0178] In the embodiment of the present application, the processor 810 also has the following functions:

[0179] Performing semantic analysis on a first text corresponding to the first audio to determine a scene type corresponding to the first audio;

[0180] Target scene effect information matching the scene type is acquired, where the target scene effect information includes target scene sound effects and / or visualized target multimedia content.

[0181] In the embodiment of the present application, the processor 810 also has the following functions:

[0182] Obtain a second text corresponding to a second audio before the first audio;

[0183] According to the second text, a semantic analysis is performed on the first text corresponding to the first audio to determine the scene type corresponding to the first audio.

[0184] In the embodiment of the present application, the processor 810 also has the following functions:

[0185] The first audio is converted into text to obtain a first text corresponding to the first audio.

[0186] In the embodiment of the present application, the target scene effect information includes the target scene sound effect, and the processor 810 also has the following functions:

[0187] Mixing the first audio with the target scene sound effect to obtain the target audio;

[0188] Output the target audio.

[0189] In the embodiment of the present application, the target scene effect information includes the target scene sound effect, and the processor 810 also has the following functions:

[0190] Get the scene sound effect library that matches the scene type;

[0191] Get the target scene sound effect from the scene sound effect library.

[0192] In the embodiment of the present application, the processor 810 also has the following functions:

[0193] Get user information;

[0194] Get the target scene sound effect that matches the user information from the scene sound effect library.

[0195] In the embodiment of the present application, the target scene effect information includes visualized target multimedia content, and the processor 810 also has the following functions:

[0196] Output visual target multimedia content.

[0197] In the embodiment of the present application, the processor 810 also has the following functions:

[0198] Detection effect adding operation;

[0199] In response to the effect adding operation, a semantic analysis is performed on the first text corresponding to the first audio to determine the scene type corresponding to the first audio.

[0200] The embodiment of the present application discloses a computer-readable storage medium on which executable program codes are stored. When the executable program codes are executed by a processor, the method described in the electronic device in the embodiment of the present application is implemented.

[0201] The embodiment of the present application discloses a computer program product. When the computer program product is executed on a computer, the computer implements the method described in the electronic device in the embodiment of the present application.

[0202] An embodiment of the present application discloses an application publishing platform, which is used to publish a computer program product. When the computer program product runs on a computer, the computer implements the method described in the electronic device in the embodiment of the present application.

[0203] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0204] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in one embodiment" or "in some embodiments" appearing throughout the specification may not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments. The above description of each embodiment tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. For the sake of brevity, this article will not repeat them.

[0205] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, object A and / or object B can mean: object A exists alone, object A and object B exist at the same time, and object B exists alone.

[0206] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0207] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.

[0208] The modules described above as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules; they may be located in one place or distributed on multiple network units; some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.

[0209] In addition, all functional modules in the embodiments of the present application may be integrated into one processing unit, or each module may be a separate unit, or two or more modules may be integrated into one unit; the above-mentioned integrated modules may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0210] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.

[0211] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0212] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0213] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0214] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0215] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for obtaining scene effect information, characterized in that: include: Performing semantic analysis on a first text corresponding to a first audio to determine a scene type corresponding to the first audio; Target scene effect information matching the scene type is acquired, where the target scene effect information includes target scene sound effects and / or visualized target multimedia content.

2. The method according to claim 1, characterized in that The performing semantic analysis on the first text corresponding to the first audio to determine the scene type corresponding to the first audio includes: Obtain a second text corresponding to a second audio before the first audio; According to the second text, a semantic analysis is performed on the first text corresponding to the first audio to determine the scene type corresponding to the first audio.

3. The method according to claim 1, characterized in that Before performing semantic analysis on the first text corresponding to the first audio to determine the scene type corresponding to the first audio, the method further includes: The first audio is converted into text to obtain a first text corresponding to the first audio.

4. The method according to claim 1, characterized in that The target scene effect information includes the target scene sound effect. After acquiring the target scene effect information matching the scene type, the method further includes: Mixing the first audio with the target scene sound effect to obtain target audio; The target audio is output.

5. The method according to any one of claims 1 to 4, characterized in that: The target scene effect information includes a target scene sound effect, and obtaining the target scene sound effect matching the scene type includes: Acquire a scene sound effect library matching the scene type; Obtain target scene sound effects from the scene sound effects library.

6. The method according to claim 5, characterized in that The step of acquiring a target scene sound effect from the scene sound effect library comprises: Get user information; From the scene sound effect library, a target scene sound effect matching the user information is obtained.

7. The method according to any one of claims 1 to 4, characterized in that: The target scene effect information includes the visualized target multimedia content. After acquiring the target scene effect information matching the scene type, the method further includes: The visualized target multimedia content is output.

8. The method according to any one of claims 1 to 4, characterized in that: The performing semantic analysis on the first text corresponding to the first audio to determine the scene type corresponding to the first audio includes: Detection effect adding operation; In response to the effect adding operation, a semantic analysis is performed on a first text corresponding to the first audio to determine a scene type corresponding to the first audio.

9. A device for acquiring scene effect information, characterized in that: include: A scene type identification unit, configured to perform semantic analysis on a first text corresponding to a first audio to determine a scene type corresponding to the first audio; The effect information acquisition unit is used to acquire target scene effect information matching the scene type, wherein the target scene effect information includes target scene sound effects and / or visualized target multimedia content.

10. An electronic device, characterized in that: include: A memory storing executable program code; and a processor coupled to the memory; The processor calls the executable program code stored in the memory, and when the executable program code is executed by the processor, the processor implements the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having executable program code stored thereon, characterized in that: When the executable program code is executed by a processor, the method according to any one of claims 1 to 8 is implemented.