Information display method, electronic equipment, storage medium and product

The AR method enhances user experience by processing audio and image data to provide personalized information overlays on multimedia content, addressing the limitations of traditional broadcast and AR technologies in dynamic content identification.

CN120321469APending Publication Date: 2025-07-15BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510540174.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Traditional broadcast television's one-way transmission model struggles to meet viewers' growing demands for information diversity and personalization, and existing AR technologies face challenges in accurately identifying dynamic multimedia content for providing additional information.

Method used

An AR-based method that utilizes audio and image data processing to identify and retrieve relevant information associated with multimedia content, enabling real-time enhancement of user experience by overlaying additional information on the display.

Benefits of technology

Enables users to access diverse and personalized information without the need for additional search tools, improving the accuracy and responsiveness of AR interactions with dynamic content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321469A_ABST
    Figure CN120321469A_ABST
Patent Text Reader

Abstract

The invention provides an information display method, electronic equipment, a storage medium and a product. Specifically, a server can obtain image data and audio data collected and sent by a first terminal; according to the audio data and at least one stored audio sample, determining a target audio sample corresponding to the audio data and a first time node corresponding to the audio data in the target audio sample; determining at least one second associated object corresponding to the first time node in the first associated objects according to the first time node; determining a target associated object according to the at least one second associated object and the image data; the display information corresponding to the target associated object is obtained, and the display information is sent to the first terminal, so that the first terminal can display the display information of the target associated object, and the diversified and personalized requirements of a user for information are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of augmented reality technology, and in particular, to a method for displaying information, an electronic device, a storage medium, and a product. Background Art

[0002] With the rapid rise of the Internet and new media platforms, the viewing habits of audiences have changed greatly. The one-way transmission mode of traditional radio and television only provides limited information and gradually fails to meet the diverse and personalized information needs of audiences. Summary of the Invention

[0003] In view of this, the purpose of the present disclosure is to provide a method for displaying information, an electronic device, a storage medium, and a product.

[0004] Based on the above purpose, the present disclosure provides a method for displaying information, which is applied to a server; the server stores at least one audio sample and a corresponding first associated object; wherein, the audio sample includes a time stamp;

[0005] The method includes:

[0006] Obtain image data and audio data collected and sent by a first terminal;

[0007] According to the audio data and the at least one audio sample, determine a target audio sample corresponding to the audio data and a first time node corresponding to the audio data in the target audio sample;

[0008] According to the first time node, screen at least one second associated object corresponding to the first time node from the first associated object corresponding to the target audio sample;

[0009] Determine a target associated object according to the at least one second associated object and the image data;

[0010] Obtain display information corresponding to the target associated object and send the display information to the first terminal.

[0011] Based on the same inventive concept, an embodiment of the present disclosure further provides a method for displaying information, which is applied to a first terminal; the method includes:

[0012] In response to obtaining a first operation of a user, display a first interactive interface; wherein, the first interactive interface includes a scanning option;

[0013] In response to obtaining a second operation of the user for the scanning option, scan a display screen of a second terminal to obtain image data and collect ambient sound to obtain audio data; wherein, the ambient sound includes sound emitted by the second terminal;

[0014] Send the image data and the audio data to a server, so that the server determines a target associated object in the image data based on the image data and the audio data;

[0015] Obtain display information corresponding to the target associated object through the server, and display the display information.

[0016] Based on the same inventive concept, an embodiment of the present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for displaying information as described in any one of the above is implemented.

[0017] Based on the same inventive concept, an embodiment of the present disclosure further provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method for displaying information as described in any one of the above.

[0018] Based on the same inventive concept, an embodiment of the present disclosure further provides a computer program product, including computer program instructions. When the computer program instructions run on a computer, the computer is caused to execute the method for displaying information as described in any one of the above.

[0019] As can be seen from the above, the present disclosure provides a method, an electronic device, a storage medium, and a product for displaying information. Specifically, the server can obtain image data and audio data collected and sent by a first terminal; determine a target audio sample corresponding to the audio data and a first time node corresponding to the audio data in the target audio sample according to the audio data and at least one stored audio sample; screen at least one second associated object corresponding to the first time node from the first associated object corresponding to the target audio sample according to the first time node; determine a target associated object according to the at least one second associated object and the image data; obtain display information corresponding to the target associated object, and send the display information to the first terminal, so that the first terminal can display the display information of the target associated object, meeting the user's diverse and personalized information needs. Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1Shows a schematic diagram of a scenario in the related art;

[0022] Figure 2A Shows a schematic diagram of an application scenario of a method for displaying information provided by an embodiment of the present disclosure;

[0023] Figure 2B Shows a partial architecture schematic diagram of a system for displaying information provided by an embodiment of the present disclosure;

[0024] Figure 3A Shows a schematic diagram of a display interface of a first terminal provided by an embodiment of the present disclosure;

[0025] Figure 3B Shows a schematic diagram of a display interface of another first terminal provided by an embodiment of the present disclosure;

[0026] Figure 3C Shows a schematic diagram of a display interface of another first terminal provided by an embodiment of the present disclosure;

[0027] Figure 3D Shows a schematic diagram of a display interface of another first terminal provided by an embodiment of the present disclosure;

[0028] Figure 4 Shows a schematic flowchart of a method for displaying information provided by an embodiment of the present disclosure;

[0029] Figure 5 Shows a schematic flowchart of another method for displaying information provided by an embodiment of the present disclosure;

[0030] Figure 6 Shows a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0031] To make the objectives, technical solutions and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.

[0032] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the field to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or items appearing before the word cover the elements or items listed after the word and their equivalents, without excluding other elements or items. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0033] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0034] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that executes the operations of the technical solutions of the present disclosure according to the prompt message.

[0035] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0036] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manners of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manners of the present disclosure.

[0037] Figure 1Shows a schematic diagram of a scenario in the related art. Scenario 100 includes a display device 101 and a user 102. The display device 101 can be a television, mobile phone, tablet computer, medical service terminal, advertising screen, projector, etc. The display device 101 can play pre-recorded multimedia information, such as TV programs, advertisements, documentaries, movies and TV dramas, etc. The user 102 watches the multimedia information played by the display device 101. Since the multimedia information is pre-recorded, the user 102 can only obtain fixed information. When the user 102 is interested in some objects (such as goods, artworks, cultural relics) in the multimedia information and wants to know more information, the user 102 needs to use a search tool (such as a search engine, a large language model for search, a shopping app, etc.) to search, with poor convenience and insufficient pertinence.

[0038] Augmented Reality (AR) technology is a technology that combines virtual information with the real world. By superimposing virtual content such as images, videos, audio, or 3D models generated by a computer onto the real scene seen by the user, it enhances the user's perception and interaction experience of the real world.

[0039] Based on this, the inventors of the present disclosure attempt to apply AR technology to Scenario 100, so that the user 102 can understand more detailed information about the objects of interest without using a search tool while watching the multimedia information of the display device 101.

[0040] The inventors further found that ordinary AR scanning relies on static image recognition. When the user 102 performs AR scanning on the playback screen of the display device 101, it is difficult to match the objects in the dynamically played multimedia information (such as video) due to reasons such as viewing angle and light changes.

[0041] In view of this, the embodiments of the present disclosure provide a method, an electronic device, a storage medium, and a product for displaying information. Specifically, the server can obtain the image data and audio data collected and sent by the first terminal; determine the target audio sample corresponding to the audio data and the first time node corresponding to the audio data in the target audio sample according to the audio data and at least one stored audio sample; screen at least one second associated object corresponding to the first time node from the first associated objects corresponding to the target audio sample according to the first time node; determine the target associated object according to the at least one second associated object and the image data; obtain the display information corresponding to the target associated object, and send the display information to the first terminal, so that the first terminal can display the display information of the target associated object, meeting the user's diverse and personalized needs for information.

[0042] To make the technical solutions of the present disclosure clearer and easier to understand, the following introduces the scenario architecture of the method for displaying information provided in the embodiments of the present disclosure with reference to the accompanying drawings.

[0043] Figure 2A A schematic diagram of an application scenario of a method for displaying information provided in an embodiment of the present disclosure is shown; Figure 2B A partial architecture schematic diagram of a system for displaying information provided in an embodiment of the present disclosure is shown. The method for displaying information provided in the embodiments of the present disclosure includes, but is not limited to, being applied to an application scenario such as Figure 2A shown. As Figure 2A shown, this application scenario includes a first terminal 202, a server 204, and a third-party platform 206. Among them, the first terminal 202, the server 204, and the third-party service interface 206 can all be connected through a wired or wireless communication network. The first terminal 202 includes, but is not limited to, a mobile phone, a mobile computer, a tablet computer, a smart wearable device, a personal digital assistant (PDA), or other electronic devices capable of implementing the above functions. The server 204 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0044] Further, as Figure 2B shown, the first terminal 202 includes a sound collection module 2021, an image collection module 2023, a rendering engine 2025, and an interaction interface 2027. Here, AR software is configured in the first terminal 202, including but not limited to AR applets, AR applications, etc. The rendering engine 2025 and the interaction interface 2027 can be components of the AR software, or the rendering engine 2025 and the interaction interface 2027 can be called through the AR software. The sound collection module 2021 is used to collect ambient sound to obtain audio data. The image collection module 2023 is used to collect image data including the display screen of the display device. The audio data and the image data can be sent to the server 204. The rendering engine 2025 is used to superimpose enhanced information, such as a 3D model, a text floating layer, etc., on the display screen. Exemplarily, the enhanced information can come from the server 204. The interaction interface 2027 is used to obtain the operation actions of the user 102 and support operations such as clicking and swiping.

[0045] The server 204 includes a program database 2041, an object database 2043, and an identification engine 2045. The program database 2041, the object database 2043, and the identification engine 2045 can communicate with each other through an internal Application Programming Interface (API).

[0046] The program database 2041 can store multiple program metadata, and each program metadata corresponds to a program. It should be noted that the programs here are not limited to TV programs, but can also be advertising films, documentaries, promotional videos, movies, etc. In other words, any pre-recorded multimedia information can be stored in the program database 2041, and the present disclosure does not limit this. It should be understood that for programs with a long duration, they can be split into multiple program segments for storage respectively, and the present disclosure does not limit this.

[0047] In some embodiments, the program metadata includes a program identifier (ID), a program description, an audio sample, and an associated object list.

[0048] The program description can record information such as the theme and recording scene of the program.

[0049] The audio sample includes at least one first audio feature group and a time stamp corresponding to each first audio feature group. It should be noted that the first audio feature group can include multiple frames of first audio features in chronological order, such as 3 frames, 5 frames, 6 frames, etc., and the present disclosure does not limit this.

[0050] The first audio feature is obtained based on the program audio according to a predetermined audio feature extraction method. Here, the predetermined audio feature extraction method is the same as the following audio data feature extraction method for facilitating the matching of audio data. Grouping the first audio features corresponding to each frame of the program audio in order, with a predetermined number of frames in a group, multiple first audio feature groups can be obtained. Exemplarily, if 100 frames of first audio features are extracted from an audio program and divided into groups of 5 frames each, 20 first audio feature groups can be obtained. For example, the first frame, the second frame, the third frame, the fourth frame, and the fifth frame are the first group; the sixth frame, the seventh frame, the eighth frame, the ninth frame, and the tenth frame are the second group... the ninety-sixth frame, the ninety-seventh frame, the ninety-eighth frame, the ninety-ninth frame, and the one-hundredth frame are the twentieth group.

[0051] The time stamp corresponding to each first audio feature group may be the time stamp of one of the frames, such as the first frame (e.g., the 1st frame of the first group), the middle frame (e.g., the 3rd frame of the first group), or the last frame (e.g., the 5th frame of the first group). The time stamp corresponding to each first audio feature group may also be formed based on the time stamps of each frame within the group, such as the average value of the time stamps of each frame. For the confirmation method of the time stamp, it can be flexibly set according to needs, and the present disclosure does not limit this.

[0052] The associated object list includes at least one associated object and at least one time stamp corresponding to the associated object. Here, the associated object may be an object that appears in the program, such as a water cup, wheat, a car, a signboard, etc. The associated object may also be a building that appears in the program, such as a high-rise building, a museum, etc. The associated object may also be a person who appears in the program, such as a host, an audience member, an actor, etc. In other words, any object that can be recognized by image recognition in the program can be used as an associated object, and the present disclosure does not limit this. It should be understood that in a program, an associated object may appear only once or may appear multiple times, such as at the 1st minute, the 5th minute, and 11 minutes and 25 seconds. Thus, an associated object may correspond to one time stamp or may correspond to multiple time stamps.

[0053] Based on this, after the program database 2041 receives the query request of the recognition engine 2045, it can filter out at least one candidate associated object from the associated object list according to the time information in the query request, so as to facilitate the recognition engine 2045 to accurately recognize the image data.

[0054] The object database 2043 stores at least one object display information. It should be noted that each object display information corresponds to an associated object. Specifically, the object display information includes a three-dimensional modeling file (such as GL Binary) of the associated object, attribute information, etc. It should be understood that for different associated objects, their attribute information may be different. For example, if the associated object is a museum, the corresponding attribute information includes but is not limited to product parameters (such as ticket price, business hours, ticket purchase link, etc.), cultural background, exhibit features, etc. Another example is that if the associated object is wheat, the corresponding attribute information includes but is not limited to product parameters (such as purchase link), variety characteristics, production place, etc. Another example is that if the associated object is an actor, the corresponding attribute information includes but is not limited to the debut time, main works, main work viewing links, etc.

[0055] The recognition engine 2045 can extract the features of the audio data and compare them with the audio samples in the program database to determine the corresponding program ID and accurate time stamp. The recognition engine 2045 can also filter out candidate associated objects from the object database based on the time stamp and perform candidate associated object detection on the image data to determine the target associated object.

[0056] The third - party service interface 206 may include an advertising platform, an e - commerce interface, etc. By using the third - party service interface 206, corresponding services can be provided to the user 206.

[0057] It should be noted that the above - mentioned application scenarios and architectures are only shown for the convenience of understanding the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not restricted in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenarios and architectures.

[0058] Next, a method for displaying information provided by the embodiments of the present disclosure will be described in detail.

[0059] See Figure 4 The flowchart of the method for displaying information shown, the method includes:

[0060] First, the first terminal 202 executes the following steps S401 - S403:

[0061] S401: As Figure 3A shown, in response to obtaining a first operation of the user 102, display a first interaction interface 301; wherein, the first interaction interface 301 includes a scan option 302. It should be noted that the first operation may be that the user 102 clicks on the icon of the AR mini - program, and the present disclosure does not limit this.

[0062] S402: In response to obtaining a second operation of the user 102 for the scan option 302, scan the display screen of the second terminal (corresponding to the display device 101) to obtain image data, and collect ambient sound to obtain audio data; wherein, the ambient sound includes the sound emitted by the second terminal. Here, the image data is environmental data including the display screen of the second terminal (corresponding to the display device 101).

[0063] It should be noted that the ambient sound refers to the sound in the space where the first terminal 202 is located. Therefore, the ambient sound includes the sound generated by the multimedia information (such as a program) played by the display device 101, and may also include the sounds made by people and objects in the space, such as the low - frequency noise generated by the operation of the display device 101, the sound made by the user 102. In addition, the ambient sound may also include the sound coming in through the window into this space. Thus, it can be seen that the ambient sound is not a single sound, but a mixture of sounds from multiple sources.

[0064] In some embodiments, the ambient sound is collected by the sound collection module 2021 to obtain audio data.

[0065] In some embodiments, the image acquisition module 2023 scans the display screen of the second terminal (corresponding to the display device 101) to obtain image data. It should be noted that the image acquisition module 2023 can scan the entire display screen of the display device 101, or can scan a part of the display screen (such as the lower left corner, upper right corner, etc. of the display screen).

[0066] S403: Send the image data and audio data to the server 204 so that the server 204 determines the target associated object in the image data based on the image data and audio data.

[0067] Next, the server 204 executes the following steps S404 to S408:

[0068] S404: Obtain the image data and audio data collected and sent by the first terminal 202.

[0069] S405: Determine the target audio sample corresponding to the audio data and the first time node corresponding to the audio data in the target audio sample according to the audio data and at least one audio sample.

[0070] In some embodiments, S405 specifically includes:

[0071] S4051: Frame and preprocess the audio data to obtain multiple frames of second audio with a time sequence.

[0072] Exemplarily, first, based on a preset time rule, the audio data is framed to obtain at least one frame of third audio; among them, at least one frame of third audio has a time sequence.

[0073] Here, the preset time rule can be 25 ms per frame with 10 ms overlapping framing. That is to say, the audio data is divided into multiple segments, and the duration of each segment is 25 ms. There is a 10 ms overlap between each segment and the adjacent previous segment. For example, the first frame is the audio segment from 0 to 25 ms in the audio data, the second frame is the audio segment from 15 to 40 ms in the audio data, the third frame is the audio segment from 30 to 55 ms in the audio data, and so on. Those skilled in the art can understand that any two adjacent frames have a chronological order in time.

[0074] Sampling in this way can avoid missing the changes in the sound.

[0075] It should be understood that the preset time rule can be modified as needed, for example, adjusted to 50 ms per frame with 20 ms overlapping framing. The present disclosure does not limit this.

[0076] In audio data, there may be silent segments. For example, when the host pauses during a program, such audio segments can be directly removed, which helps save computing power and avoid waste of computing power.

[0077] Next, based on a preset energy threshold, at least one frame of the third audio is screened to obtain at least one fourth audio.

[0078] Here, the energy threshold can refer to a critical value of the amplitude (or power) of the audio. When the amplitude of the audio signal is lower than this threshold, this frame of audio can be considered "silent". The amplitude can be expressed in decibels (dB), or it can be expressed as the absolute value of the amplitude (such as a floating point number between 0 and 1).

[0079] Exemplarily, the energy threshold can be -60 dB. When the amplitude of any frame of the third audio is lower than -60 dB, this frame of the third audio can be considered silent and can be directly removed. That is, the third audio with an amplitude exceeding the energy threshold is screened out as the fourth audio.

[0080] Finally, pre-emphasis and high-pass filtering are performed on each frame of the fourth audio to obtain multiple frames of the second audio.

[0081] Pre-emphasis is an audio processing technique mainly used to enhance the high-frequency components of the audio signal, similar to "turning up the volume". For example, when a person speaks, the high-frequency sounds are prone to attenuation. By pre-emphasizing the high frequencies, the voice can be made clearer. Here, a first-order high-pass filter can be used to implement pre-emphasis, and the coefficient of its transfer function can be 0.97. The present disclosure does not limit the specific implementation manner of pre-emphasis.

[0082] High-pass filtering is a signal processing technique used to remove the components of the audio signal below a specific frequency (referred to as the cut-off frequency), while retaining the signal components above that frequency. Here, high-pass filtering can remove the low-frequency noise in the fourth audio, such as the humming of the air conditioner and the working noise of the display device 101, and only retain the useful mid-high frequencies such as human voices and music. The present disclosure does not limit the specific implementation manner of high-pass filtering.

[0083] S4052: Feature extraction is performed on each frame of the second audio to obtain the corresponding second audio features. S4052 can transform the second audio into a set of digital features, which is convenient for determining whether two sounds are the same.

[0084] Exemplarily, the first audio feature is Mel-Frequency Cepstral Coefficients (MFCC), which is an audio feature extracted based on a Mel filter bank. Here, the second audio feature is also MFCC, and the first audio feature and the second audio feature are extracted in the same way.

[0085] An embodiment of the present disclosure provides a specific example to exemplarily illustrate the extraction of Mel-Frequency Cepstral Coefficients:

[0086] First, perform a Fast Fourier Transform (FFT) on the second audio to obtain a spectrogram; here, the Fast Fourier Transform is an efficient algorithm that can be used to calculate the Discrete Fourier Transform (DFT) and its inverse transform. With the help of FFT, the second audio can be converted from a waveform into a spectrogram, showing the volume levels of different frequencies.

[0087] Next, pass the spectrogram through a Mel Filter Bank (MFB) to obtain a Mel spectrogram. The human ear has different sensitivities to low frequencies (such as drum sounds) and high frequencies (such as bird calls). The Mel filter bank can simulate the human ear with virtual ears and convert the spectrogram into "Mel frequencies" that are more in line with human hearing. Optionally, the number of channels of the Mel filter bank can be 40, 50, etc., and each channel can correspond to a virtual ear. The present disclosure does not limit this.

[0088] Finally, take the logarithm of the Mel spectrogram and perform a Discrete Cosine Transform (DCT), and retain the predetermined number of coefficients in the front as the second audio feature. Here, the coefficients of the Discrete Cosine Transform have a sequential order, and the predetermined number in the front refers to the predetermined number in the front order, such as 13, 15, etc., and the present disclosure does not limit this.

[0089] It should be noted that the loudness of the sound heard by the human ear is not linear (for example, when the volume doubles, the human perception may only increase a little bit). Taking the logarithm can compress the data into a mode that is closer to human perception.

[0090] S4051 to S4052 can be executed by the recognition engine 2045 of the server 204.

[0091] S40531: In response to determining that the multi-frame second audio features match any one of the first audio feature groups in any audio sample, determine the corresponding audio sample as the target audio sample.

[0092] Here, each program can correspond to an audio sample, and different audio samples can be distinguished using the program ID.

[0093] The audio sample may include multiple different first audio feature groups, which are not limited in the present disclosure.

[0094] It should be noted that before performing the matching step, the recognition engine 2045 may call the audio samples in the program database 2041, which is not limited in the present disclosure.

[0095] For the specific steps of determining that the multi-frame second audio feature matches any first audio feature group in any audio sample, examples are as follows:

[0096] First, calculate the minimum cumulative distance between the multi-frame second audio feature and each first audio feature group in each of the audio samples respectively. It should be noted that the minimum cumulative distance can be implemented by using the Dynamic Time Warping (DTW) algorithm, which specifically includes constructing a cumulative distance matrix and finding the shortest path. The present disclosure does not elaborate on this.

[0097] In this way, regardless of whether the number of frames of the multi-frame second audio feature is the same as the number of frames of the first audio feature group, it can accurately determine whether the two are similar, achieving the technical effect of anti-time shift. For example, the audio data recorded by the user 102 is short in time, and the corresponding number of frames of the second audio feature is small. In this way, it can also be effectively aligned with the first audio feature group.

[0098] Next, in response to determining that the minimum cumulative distance between the multi-frame second audio feature and any of the first audio feature groups is less than a preset threshold, the multi-frame second audio feature matches the first audio feature group.

[0099] Exemplarily, the preset threshold may be 5. If the minimum cumulative distance between the B-th first audio feature group of the A-th audio sample and the multi-frame second audio feature is 4, since 4 is less than 5, the A-th audio sample is the target audio sample, and the B-th first audio feature group is successfully matched.

[0100] S4054: Determine the first time node corresponding to the audio data in the target audio sample according to the time mark corresponding to the matched first audio feature group.

[0101] In some embodiments, S4054 includes:

[0102] Based on the time mark (such as a time stamp) corresponding to the first audio feature group, determine a second time node; here, if the time stamp corresponding to the first audio feature group is 5 min, then the second time node is 5 min.

[0103] Considering network latency, the second time node is adjusted by adjusting the rules to avoid affecting the screening of associated objects due to time deviation.

[0104] Based on the preset adjustment rules, the first time node is obtained by adjusting the second time node.

[0105] Exemplarily, the adjustment rule is the second time node ± 5s. If the second time node is 5 minutes, the first time node can be from 4 minutes 55 seconds to 5 minutes 05 seconds.

[0106] S406: According to the first time node, at least one second associated object corresponding to the first time node is screened out from the first associated objects corresponding to the target audio sample.

[0107] Here, the first associated object refers to all the associated objects in the associated object list corresponding to the program ID of the target audio sample. Since each associated object in the associated object list corresponds to at least one timestamp, if any timestamp of any first associated object falls within the range of the first time node, the first associated object is screened as a second associated object.

[0108] Obviously, by determining the first time node from the audio data and using the first time node to screen the first associated objects to obtain the second associated objects, it helps to reduce the recognition difficulty of subsequent image data, improve the recognition effect, and reduce the computing power consumption.

[0109] S407: Determine the target associated object according to the at least one second associated object and the image data.

[0110] In some embodiments, determining the target associated object specifically includes:

[0111] S4071: Perform a first image detection on the image data to obtain at least one candidate image corresponding to the at least one second associated object.

[0112] Optionally, the performing a first image detection on the image data to obtain at least one candidate image corresponding to the at least one second associated object specifically includes:

[0113] First, determine the number of the at least one second associated object and the category of each second associated object. For example, the number of the second associated objects is 3, which are a water cup, a chair, and a table lamp respectively, then the corresponding categories of the second associated objects are a water cup, a chair, and a table lamp.

[0114] It should be noted that the category of the second associated object can be directly obtained from the associated object list.

[0115] Next, load a first image processing model that matches the quantity and the categories of each of the second associated objects.

[0116] Here, the first image processing model can be the dynamic YOLOv6-tiny model. The YOLOv6-tiny model can recognize images of 320×320 pixels, with fast recognition speed and low computing power consumption.

[0117] It should be understood that the first image processing model can be dynamically loaded. If there are N second associated objects, only N + 1 nodes of the model output layer are retained (N target categories + 1 background category). This way can reduce the amount of calculation and improve the recognition speed. Exemplarily, if the second associated objects are 5 types of objects, the YOLOv6-tiny model only retains 5 + 1 output options, instead of outputting the full amount, such as 1000 types.

[0118] It should be noted that the first image processing model is updated using federated learning. Exemplarily, if an object is recognized as failed multiple times, the model will collect error cases and adjust the model parameters.

[0119] Adopting such a technical solution, the recognition speed can be increased by 40%, for example, reduced from 100 ms to 60 ms.

[0120] Finally, use the first image processing model to detect the image data to obtain at least one candidate object image. By using the first image processing model, the potential candidate object image can be quickly located.

[0121] S4072: Perform a second image detection on the candidate image to obtain the target associated object. Optionally, the second image detection can adopt a Convolutional Neural Network (CNN), such as MobileNetV2.

[0122] Thus, through the first image detection, the rough detection of the image data is realized, and through the second image detection, the fine detection of the image data is realized. With the help of two-stage detection, the detection efficiency is improved while the detection accuracy is enhanced.

[0123] S408: Obtain the display information corresponding to the target associated object and send the display information to the first terminal 202. Here, the recognition engine 2045 can call the object database 2043 to execute S408.

[0124] In some embodiments, as shown above, the display information includes a 3D modeling file; the obtaining the display information corresponding to the target associated object specifically includes:

[0125] Obtain and determine the level of the first terminal according to the device information of the first terminal.

[0126] Optionally, when the first terminal 202 communicates with the server 204 for the first time, device information can be sent to the server 204. Here, the device information can include the floating-point performance of the graphics processing unit (GPU).

[0127] Exemplarily, if the floating-point performance < 1 TFLOPS, it is determined that the first terminal is a low-end device; if the floating-point performance ≥ 1 TFLOPS, it is determined that the first terminal is a high-end device.

[0128] Based on the level, obtain a 3D modeling file corresponding to the target associated object and matching the level.

[0129] It should be noted that the 3D modeling file can be of the first type and the second type. The first type is used to form a reduced version of the 3D model (corresponding to the number of patches < 5k, texture resolution 512×512). The second type is used to form a high-definition 3D model (corresponding to the number of patches 50k, 4K texture + PBR material).

[0130] If the first terminal 202 is a low-end device, it corresponds to the 3D modeling file of the first type; if the first terminal 202 is a high-end device, it corresponds to the 3D modeling file of the second type.

[0131] Next, the first terminal 202 executes steps S409 to S4

[0132] S409: Obtain the display information corresponding to the target associated object through the server 204 and display the display information.

[0133] In some embodiments, as Figure 3B shown, the displaying of the display information specifically includes:

[0134] Render a 3D model according to the 3D modeling file; it should be noted that the rendering of the 3D model can be executed by the rendering engine 2025. If the rendering frame rate is lower than 20 FPS, the rendering of the 3D model can automatically switch to the reduced version of the 3D model.

[0135] Here, if the performance of the first terminal 202 is insufficient, the cloud GPU can be called to render the 3D model and push it to the first terminal 202 in the form of a video stream.

[0136] Display the 3D model 3041 on the second interaction interface 303.

[0137] Optionally, the second interaction interface 303 includes a floating layer 304. The 3D model 3041 is displayed on the floating layer 304. The base layer of the second interaction interface 303 can display an environmental image including the display screen of the first terminal 101.

[0138] It should be noted that the floating layer 304 can also display the description information 3042 (corresponding to the information other than the product link in the attribute information) and the product link 3043. It should be understood that when the user 102 clicks on the product link 3043, the third-party service interface 206 can be called through the server 204.

[0139] Exemplarily, referring to Figure 3B , the 3D model 3041 is a 3D model of a tourist town. The description information 3042 includes an introduction to the tourist town. The product link 3043 may include a connection for purchasing tickets for the town.

[0140] It should be understood that the user 102 can select the display information on the floating layer 304, for example, only display the 3D model 3041, or display both the 3D model 3041 and the description information 3042 at the same time.

[0141] In addition, the first terminal 202 also supports hot switching technology and can asynchronously load high-definition 3D models in the background. When the user 102 clicks on the "Enhanced Image Quality" option of the second interaction interface 303, seamless replacement can be achieved.

[0142] In some embodiments, as Figure 3C shown, the second interaction interface 303 includes an interaction component 3044; the solution further includes:

[0143] In response to obtaining a third operation of the user 102 on the interaction component 3044, controlling the 3D model 3041 to perform a corresponding action.

[0144] Exemplarily, if the user 102 clicks on the "Rotate Left" in the interaction component 3044, then control the 3D model 3041 to rotate left.

[0145] In addition, the method further includes: in response to determining that the audio data fails to correspond to any of the audio samples, the server 204 determines a target association object according to the image data and each of the first association objects.

[0146] It should be noted that compared with S407, the number of first association objects is significantly more than that of second association objects, and a complete first image processing model needs to be called, resulting in relatively high computing power consumption. Here, the first association object is not limited to a certain program, but all association objects in the program database.

[0147] Optionally, in response to determining that the audio data fails to correspond to any of the audio samples, the server 204 may send a failure message to the first terminal 202. As Figure 3D shown, after receiving the failure message, the first terminal 202 can prompt a prompt box 305 on the third interaction interface to remind the user 102 to align the image acquisition module with the object center.

[0148] Through technologies such as image recognition algorithms and audio matching algorithms, improve the accuracy and response speed of AR recognition in dynamic display screens. It can not only perform object recognition, but also provide users with more interactive experience methods, realizing seamless interaction between the display screen and object information, which helps to expand the monetization scenarios of TV e-commerce and advertising, such as "AR scan to jump to place an order and purchase", etc.

[0149] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server, etc. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0150] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or consecutive order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0151] Based on the same inventive concept, embodiments of the present disclosure provide a method for presenting information. As Figure 5 shown, the method is applied to the server 204; the server 204 stores at least one audio sample and the corresponding first associated object, here, as Figure 2B shown, it can be stored in the program database 2041; wherein, the audio sample includes a time marker, such as a timestamp;

[0152] The method includes:

[0153] S501: As Figure 2A shown, obtain the image data and audio data collected and sent by the first terminal 202; here, both the image data and the audio data can be environmental data, as described above, and will not be elaborated further.

[0154] S503: According to the audio data and the at least one audio sample, determine the target audio sample corresponding to the audio data and the first time node corresponding to the audio data in the target audio sample; it should be noted that step S503 can refer to S405 and will not be elaborated further.

[0155] S505: According to the first time node, screen at least one second associated object corresponding to the first time node from the first associated objects corresponding to the target audio sample; it should be noted that step S505 can refer to S406 and will not be elaborated here.

[0156] S507: Determine the target associated object according to the at least one second associated object and the image data; it should be noted that step S507 can refer to S407 and will not be elaborated here.

[0157] S509: Obtain the display information corresponding to the target associated object and send the display information to the first terminal 202; it should be noted that step S509 can refer to S408 and will not be elaborated here.

[0158] In some embodiments, referring to S4051 - S4054, the audio sample includes at least one first audio feature group and its corresponding time stamp; the specific steps of determining the target audio sample corresponding to the audio data and the first time node corresponding to the audio data in the target audio sample according to the audio data and the at least one audio sample are as follows:

[0159] Perform frame splitting and preprocessing on the audio data to obtain multiple frames of second audio in time sequence.

[0160] Extract features from each frame of the second audio to obtain the corresponding second audio features.

[0161] In response to determining that the multiple frames of second audio features match any one of the first audio feature groups in any one of the audio samples, determine the corresponding audio sample as the target audio sample.

[0162] According to the time stamp corresponding to the matched first audio feature group, determine the first time node corresponding to the audio data in the target audio sample.

[0163] In some embodiments, referring to S4051, the specific steps of performing frame splitting and preprocessing on the audio data to obtain multiple frames of second audio in time sequence are as follows:

[0164] Based on a preset time rule, perform frame splitting on the audio data to obtain at least one frame of third audio; wherein, the at least one frame of third audio has a time sequence.

[0165] Based on a preset energy threshold, screen the at least one frame of third audio to obtain at least one frame of fourth audio.

[0166] Perform pre-emphasis and high-pass filtering on each frame of the fourth audio to obtain the multiple frames of second audio.

[0167] In some embodiments, referring to S4052, the extracting of features from each frame of the second audio to obtain corresponding second audio features specifically includes:

[0168] Performing a fast Fourier transform on the second audio to obtain a spectrogram;

[0169] Passing the spectrogram through a Mel filter bank to obtain a Mel spectrogram;

[0170] Taking the logarithm of the Mel spectrogram and performing a discrete cosine transform, and retaining a predetermined number of coefficients as the second audio features.

[0171] In some embodiments, referring to S4053, the determining that the multi-frame second audio features match any one of the first audio feature groups in any one of the audio samples specifically includes:

[0172] Respectively calculating the minimum cumulative distance between the multi-frame second audio features and each of the first audio feature groups in each of the audio samples;

[0173] In response to determining that the minimum cumulative distance between the multi-frame second audio features and any one of the first audio feature groups in any one of the audio samples is less than a preset threshold, the multi-frame second audio features match the first audio feature group.

[0174] In some embodiments, referring to S4054, the determining of the first time node corresponding to the audio data in the target audio sample according to the time stamp corresponding to the matched first audio feature group specifically includes:

[0175] Determining a second time node based on the time stamp corresponding to the first audio feature group;

[0176] Adjusting the second time node based on a preset adjustment rule to obtain the first time node.

[0177] In some embodiments, referring to S407, the determining of the target association object according to the at least one second association object and the image data specifically includes:

[0178] Performing a first image detection on the image data to obtain at least one candidate image corresponding to the at least one second association object;

[0179] Performing a second image detection on the candidate image to obtain the target association object.

[0180] In some embodiments, referring to S4071, the performing of the first image detection on the image data to obtain at least one candidate image corresponding to the at least one second association object specifically includes:

[0181] Determine the quantity of the at least one second associated object and the category of each second associated object;

[0182] Load a first image processing model that matches the quantity and the category of each second associated object;

[0183] Use the first image processing model to detect the image data to obtain at least one candidate object image.

[0184] In some embodiments, referring to S408, the display information includes a 3D modeling file; the obtaining of the display information corresponding to the target associated object specifically includes:

[0185] Obtain and determine the level of the first terminal according to the device information of the first terminal;

[0186] Based on the level, obtain a 3D modeling file corresponding to the target associated object and matching the level.

[0187] In some embodiments, it further includes:

[0188] In response to determining that the audio data fails to correspond to any of the audio samples, determine a target associated object according to the image data and each first associated object.

[0189] Based on the same inventive concept, an embodiment of the present disclosure provides a method for display information. The method is applied to the first terminal 202; the method includes:

[0190] As Figure 3A shown, in response to obtaining a first operation of a user, display a first interaction interface 301; wherein, the first interaction interface 301 includes a scan option 302;

[0191] In response to obtaining a second operation of the user for the scan option 302, scan the display screen of the second terminal to obtain image data, and collect ambient sound to obtain audio data; wherein, the ambient sound includes the sound emitted by the second terminal;

[0192] Send the image data and the audio data to the server 204, so that the server 204 determines the target associated object in the image data based on the image data and the audio data;

[0193] Obtain the display information corresponding to the target associated object through the server and display the display information.

[0194] In some embodiments, the display information includes a 3D modeling file; the displaying of the display information specifically includes:

[0195] As Figure 3BAs shown, render a 3D model according to the 3D modeling file;

[0196] Display the 3D model on the second interaction interface 303.

[0197] In some embodiments, as Figure 3C shown, the second interaction interface includes an interaction component 3044; the solution further includes:

[0198] In response to obtaining a third operation of the user on the interaction component, control the 3D model to perform a corresponding action.

[0199] The method of the above embodiment is used to implement the corresponding method for displaying information in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0200] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method for displaying information described in any one of the above embodiments.

[0201] Figure 6 Fig. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0202] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0203] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0204] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure), or can be externally connected to the device to provide corresponding functions. Among them, the input devices can include keyboards, mice, touchscreens, microphones, various sensors, etc., and the output devices can include displays, speakers, vibrators, indicator lights, etc.

[0205] The communication interface 1040 is used to connect to the communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. Among them, the communication module can achieve communication through wired means (such as USB, network cable, etc.), or can achieve communication through wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0206] The bus 1050 includes a path to transmit information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0207] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0208] The electronic device in the above embodiment is used to implement the corresponding method of displaying information in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0209] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to make the computer execute the method of displaying information as described in any of the foregoing embodiments.

[0210] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0211] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the method of displaying information as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0212] Based on the same inventive concept, corresponding to the method of displaying information described in any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of the computer to cause the computer and / or the processor to execute the method of displaying information. Corresponding to the execution subject of each step in each embodiment of the method of displaying information, the processor that executes the corresponding step can belong to the corresponding execution subject.

[0213] The computer program product of the above embodiment is used to cause the computer and / or the processor to execute the method of displaying information as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0214] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the idea of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, and they are not provided in detail for the sake of brevity.

[0215] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, well-known power / ground connections of integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In cases where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions are to be regarded as illustrative rather than restrictive.

[0216] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0217] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for displaying information, characterized in that, The method is applied to a server; the server stores at least one audio sample and a corresponding first associated object; wherein, the audio sample includes a time stamp; The method includes: Obtaining image data and audio data collected and sent by a first terminal; Determining a target audio sample corresponding to the audio data and a first time node corresponding to the audio data in the target audio sample according to the audio data and the at least one audio sample; Filtering at least one second associated object corresponding to the first time node from the first associated object corresponding to the target audio sample according to the first time node; Determining a target associated object according to the at least one second associated object and the image data; Obtaining display information corresponding to the target associated object and sending the display information to the first terminal.

2. The method according to claim 1, wherein The audio sample includes at least one first audio feature group and its corresponding time stamp; The step of determining a target audio sample corresponding to the audio data and a first time node corresponding to the audio data in the target audio sample according to the audio data and the at least one audio sample specifically includes: Framing and preprocessing the audio data to obtain multiple frames of second audio with a time sequence; Extracting features from each frame of the second audio to obtain corresponding second audio features; In response to determining that the multiple frames of second audio features match any one of the first audio feature groups in any one of the audio samples, determining the corresponding audio sample as the target audio sample; Determining a first time node corresponding to the audio data in the target audio sample according to the time stamp corresponding to the matched first audio feature group.

3. The method according to claim 2, wherein The step of framing and preprocessing the audio data to obtain multiple frames of second audio with a time sequence specifically includes: Framing the audio data based on a preset time rule to obtain at least one frame of third audio; wherein, the at least one frame of third audio has a time sequence; Filtering the at least one frame of third audio based on a preset energy threshold to obtain at least one frame of fourth audio; Performing pre-emphasis and high-pass filtering on each frame of the fourth audio to obtain the multiple frames of second audio.

4. The method according to claim 2, wherein The step of extracting features from each frame of the second audio to obtain corresponding second audio features specifically includes: Performing a fast Fourier transform on the second audio to obtain a spectrogram; Passing the spectrogram through a Mel filter bank to obtain a Mel spectrogram; Taking the logarithm of the Mel spectrogram and performing a discrete cosine transform, and retaining a predetermined number of coefficients as the second audio features.

5. The method according to claim 2, wherein The step of determining that the multiple frames of second audio features match any one of the first audio feature groups in any one of the audio samples specifically includes: Calculating the minimum cumulative distance between the multiple frames of second audio features and each first audio feature group in each audio sample respectively; In response to determining that the minimum cumulative distance between the multiple frames of second audio features and any one of the first audio feature groups in any one of the audio samples is less than a preset threshold, the multiple frames of second audio features match the first audio feature group.

6. The method according to claim 2, wherein Determining the first time node corresponding to the audio data in the target audio sample according to the time stamp corresponding to the matched first audio feature group specifically includes: Determining a second time node based on the time stamp corresponding to the first audio feature group; Adjusting the second time node according to a preset adjustment rule to obtain the first time node.

7. The method according to claim 1, wherein Determining the target associated object according to the at least one second associated object and the image data specifically includes: Performing a first image detection on the image data to obtain at least one candidate image corresponding to the at least one second associated object; Performing a second image detection on the candidate image to obtain the target associated object.

8. The method according to claim 7, wherein Performing a first image detection on the image data to obtain at least one candidate image corresponding to the at least one second associated object specifically includes: Determining the number of the at least one second associated object and the category of each second associated object; Loading a first image processing model matching the number and the category of each second associated object; Detecting the image data by using the first image processing model to obtain at least one candidate object image.

9. The method according to claim 1, wherein The display information includes a 3D modeling file; obtaining the display information corresponding to the target associated object specifically includes: Obtaining and determining the level of the first terminal according to the device information of the first terminal; Based on the level, obtaining a 3D modeling file corresponding to the target associated object and matching the level.

10. The method according to claim 1, characterized in that Further includes: In response to determining that the audio data fails to correspond to any of the audio samples, determining the target associated object according to the image data and each of the first associated objects.

11. A method for displaying information, characterized in that, The method is applied to a first terminal; the method includes: In response to obtaining a first operation of a user, displaying a first interaction interface; wherein, the first interaction interface includes a scanning option; In response to obtaining a second operation of the user for the scanning option, scanning a display screen of a second terminal to obtain image data, and collecting ambient sound to obtain audio data; wherein, the ambient sound includes the sound emitted by the second terminal; Sending the image data and the audio data to a server so that the server determines a target associated object in the image data based on the image data and the audio data; Obtaining the display information corresponding to the target associated object through the server and displaying the display information.

12. The method according to claim 11, characterized in that, The display information includes a 3D modeling file; Displaying the display information specifically includes: Rendering a 3D model according to the 3D modeling file; Displaying the 3D model on a second interaction interface.

13. The method according to claim 12, characterized in that, The second interaction interface includes interaction components; the solution further includes: In response to obtaining a third operation of the user for the interaction components, controlling the 3D model to perform a corresponding action.

14. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable by the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 10 or any one of claims 11 to 13.

15. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the method according to any one of claims 1 to 10 or any one of claims 11 to 13.

16. A computer program product, characterized in that, It includes computer program instructions which, when run on a computer, cause the computer to execute the method according to any one of claims 1 to 10 or any one of claims 11 to 13.