Searching method and device
By understanding the video content of the live broadcast room, using computer vision and machine learning technology to identify objects and establish object inverted index, the problem of difficulty in accurately retrieving associated live broadcast rooms in the existing technology is solved, and more intuitive and accurate search results are achieved, improving user experience and platform competitiveness.
Patent Information
- Application Number
- CN202510429331.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-29
AI Technical Summary
The existing search methods are difficult to accurately and comprehensively search related live broadcast rooms in content platforms, resulting in missed and mis-checked, affecting the search experience.
By understanding the video content in the live broadcast room, using computer vision and machine learning technology to identify objects, and establish an object inverted index, construct a hash table storage based on the relationship between the object and the live broadcast room, and query the target live broadcast room identification associated with the target object after receiving the client's search request.
Provide more intuitive, accurate and comprehensive search results, alleviate missed and missed detection problems, and improve search experience and platform competitiveness.
Smart Images

Figure CN120386942A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of Internet technologies, and in particular, to a search method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] In a content platform, the system can receive a user's search request, retrieve relevant objects, and return a list of object recommendations. Further, to enrich the search results, the system can also retrieve the live rooms associated with the relevant objects and return a list of live room recommendations.
[0003] However, the existing search methods still have problems of missed detection and misdetection, making it difficult to accurately and comprehensively retrieve associated live rooms, which affects the search experience.
[0004] It should be noted that the above content is not necessarily prior art and is not used to limit the patent protection scope of the present application. Summary of the Invention
[0005] Embodiments of the present application provide a search method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the above-mentioned technical problems.
[0006] One aspect of the embodiments of the present application provides a search method for a server, and the method includes: Receiving a search request sent by a client, where the search request carries an identifier of a target object; Based on the identifier of the target object, determining a target index entry from multiple index entries of an object inverted index; wherein, the object inverted index is constructed based on video contents of multiple live rooms, the video contents correspond to one or more objects, and each index entry corresponds to an object and indicates an identifier of a live room associated with the object; According to the target index entry, determining an identifier of a target live room associated with the target object; Generating a target live room list according to the identifier of the target live room, and pushing the target live room list to the client.
[0007] Optionally, the object inverted index is constructed through the following operations: Obtaining the video contents of the multiple live rooms, where the video contents include image frames; Based on the image frames of the multiple live rooms, determining the objects corresponding to the multiple live rooms; Based on the objects corresponding to the multiple live rooms, establishing the object inverted index.
[0008] Optionally, determining the objects corresponding to the multiple live rooms based on the image frames of the multiple live rooms includes: Determining the objects corresponding to the multiple live rooms based on the image frames of the multiple live rooms through visual analysis technology or a pre-trained image recognition model.
[0009] Optionally, the inverted index of objects is constructed through the following operations: Obtaining the video content of the multiple live rooms, where the video content includes audio; Determining the objects corresponding to the multiple live rooms based on the audio of the multiple live rooms; Establishing the inverted index of objects based on the objects corresponding to the multiple live rooms.
[0010] Optionally, the inverted index of objects is constructed through the following operations: Obtaining the video content of the multiple live rooms, where the video content includes image frames and audio; Determining the first objects corresponding to the multiple live rooms by performing image recognition on the image frames of the multiple live rooms; Determining the second objects corresponding to the multiple live rooms by performing audio analysis on the audio of the multiple live rooms; Determining the objects corresponding to the multiple live rooms according to the first objects and the second objects corresponding to the multiple live rooms; Establishing the inverted index of objects based on the objects corresponding to the multiple live rooms.
[0011] Optionally, determining the objects corresponding to the multiple live rooms according to the first objects and the second objects corresponding to the multiple live rooms includes: For each live room, determining whether the corresponding first object and second object match; In the case of a match, determining the first object or the second object as the object corresponding to the live room; In the case of no match, determining the confidence levels corresponding to the image frame and the audio respectively; In the case where the confidence level of the image frame is higher than that of the audio, determining the first object as the object corresponding to the live room; In the case where the confidence level of the audio is higher than that of the image frame, determining the second object as the object corresponding to the live room.
[0012] Optionally, the search method further includes: Obtaining the latest video content of the multiple live rooms, where the latest video content includes current image frames and / or current audio; Determine the objects currently corresponding to the multiple live rooms based on the current image frames and / or current audio of the multiple live rooms; Obtain the object list of each live room, where the object list includes the objects corresponding to the history of the live room, and the object list is determined based on the historical image frames and / or historical audio of the live room; Based on the object list of each live room, determine whether there are objects that appear for the first time among the objects currently corresponding to each live room; In the case where it is determined that there are objects that appear for the first time, incrementally update the object inverted index according to the objects that appear for the first time and the corresponding live rooms, and add the objects that appear for the first time to the object list of the corresponding live rooms.
[0013] Another aspect of the embodiments of the present application provides a search device, and the device includes: A receiving module, configured to receive a search request sent by a client, where the search request carries an identifier of a target object; A first determination module, configured to determine a target index entry from multiple index entries of an object inverted index based on the identifier of the target object; wherein, the object inverted index is constructed based on the video content of multiple live rooms, the video content corresponds to one or more objects, and each index entry corresponds to an object and indicates the identifier of the live room associated with the object; A second determination module, configured to determine the identifier of a target live room associated with the target object according to the target index entry; A pushing module, configured to generate a target live room list according to the identifier of the target live room and push the target live room list to the client.
[0014] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0015] Another aspect of the embodiments of the present application provides a computer-readable storage medium, where computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed by a processor, the method as described above is implemented.
[0016] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.
[0017] The embodiments of the present application adopting the above technical solutions may include the following advantages: The server obtains the video content of multiple live rooms, and the video content corresponds to one or more objects. An object inverted index is constructed based on the relationship between the objects and the live rooms. Among them, the object inverted index includes multiple index entries, and each index entry corresponds to an object and indicates the identifier of the live room associated with the object. The server receives a search request from the client, and the search request carries the identifier of the target object. Based on the identifier of the target object, the target index entry is determined from the multiple index entries of the object inverted index. According to the target index entry, the identifier of the target live room associated with the target object is determined. A target live room list is generated based on the identifier of the target live room and recommended to the client as the search result. It can be seen that the embodiments of the present application can accurately identify the objects being displayed or explained by understanding the video content of the live rooms, and establish an object inverted index based on the relationship between the objects and the live rooms. When the user searches for the target object, the associated target live room can be quickly obtained, and a target live room list is generated as the search result and pushed to the user. That is, more intuitive, accurate, and comprehensive search results can be provided, alleviating the problems of missed detection and false detection, thereby effectively improving the search experience and the platform competitiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0019] Figure 1 Schematically shows the operating environment diagram of the search method according to Embodiment 1 of the present application; Figure 2 Schematically shows the flowchart of the search method according to Embodiment 1 of the present application; Figure 3 Schematically shows the relationship between the live rooms and objects according to Embodiment 1 of the present application; Figure 4 Schematically shows the object inverted index according to Embodiment 1 of the present application; Figure 5 Schematically shows the hash table according to Embodiment 1 of the present application; Figure 6 Schematically shows the construction flowchart of the object inverted index according to Embodiment 1 of the present application; Figure 7 Schematically shows the construction flowchart of the object inverted index according to Embodiment 1 of the present application; Figure 8Schematically shows the construction flow chart of the object inverted index according to Embodiment 1 of the present application; Figure 9 Schematically shows the construction flow chart of the object inverted index according to Embodiment 1 of the present application; Figure 10 Schematically shows the construction flow chart of the object inverted index according to Embodiment 1 of the present application; Figure 11 Is an application example diagram of the search method according to Embodiment 1 of the present application; Figure 12 Schematically shows the block diagram of the search device according to Embodiment 2 of the present application; and Figure 13 Schematically shows the hardware architecture diagram of the computer device according to Embodiment 3 of the present application. Detailed implementation manners
[0020] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0021] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.
[0022] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and distinguish each step, and thus cannot be understood as a limitation to the present application.
[0023] First, provide the term explanations involved in the present application: APP (Application): Application program.
[0024] ID (Identifier): Identifier.
[0025] K-V (Key-Value): Key-value pair, a data structure used to store and organize related data.
[0026] URL (Uniform Resource Locator): Uniform Resource Locator.
[0027] Secondly, to facilitate the understanding of the technical solutions provided by the embodiments of the present application by those skilled in the art, the related technologies are described below: In a content platform, the system can receive a search request from a user, retrieve relevant objects, and return an object recommendation list. Further, to enrich the search results, the system can also retrieve the live rooms associated with the relevant objects and return a live room recommendation list.
[0028] However, the applicant has learned that the related search technologies still have problems of missed detection and misdetection, and it is difficult to accurately and comprehensively retrieve the associated live rooms, which affects the search experience.
[0029] For this reason, the embodiments of the present application provide a search technical solution. In this technical solution: through technologies such as computer vision, machine learning, and deep learning, the video content of the live room is understood to accurately identify the objects corresponding to the live room. An inverted index is established for the objects appearing in the live room and stored through a hash table. When receiving a search request carrying the identifier of the target object sent by the client, the identifier of the target live room associated with the target object can be queried in the hash table. The target live room list is returned to the client as the search result. It can be seen that this technical solution can provide more intuitive, accurate, and comprehensive search results, alleviate the problems of missed detection and misdetection, and thus effectively improve the search experience and platform competitiveness. See the following for details.
[0030] Finally, for the convenience of understanding, an exemplary operating environment is provided below.
[0031] As Figure 1 shown, the operating environment diagram includes: a service platform (server) 2, host terminals (4A, 4B,..., 4M), and viewer terminals (6A, 6B,..., 6N). In a live broadcast scenario, the host terminals (4A, 4B,..., 4M) log in to the service platform 2, and the live broadcast data is pushed to the viewer terminals (6A, 6B,..., 6N) in real time through the service platform 2.
[0032] The service platform 2 can provide search services and live room services, and it can be a single server, a server cluster, or a cloud computing service center. The server can include a business server, a video cloud server, etc.
[0033] The host terminals (4A, 4B,..., 4M) are used to generate live broadcast data in real time and perform the push operation of the live broadcast data. The live broadcast data can include audio data or video data. The host terminal can be an electronic device such as a smart phone or a tablet computer. Of course, the host terminal can be a virtual computing instance within the service platform 2.
[0034] Audience terminals (6A, 6B, …, 6N) can be configured to receive live data from the host terminal in real time. The audience terminals (6A, 6B, …, 6N) can be any type of computing device, such as a smart phone, a tablet device, a laptop computer, a smart TV, a vehicle-mounted terminal, etc. The audience terminals (6A, 6B, …, 6N) can be built with a browser or a dedicated program, and receive the live data through the browser or the dedicated program to output content to the user. The content can include video, audio, comments, text data, and / or the like.
[0035] The audience terminals (6A, 6B, …, 6N) can include a search interface and a player. Among them, the search interface can be configured with a search box, and the search box can be used to receive query text input by the user. The audience terminal can generate a search request according to the received query text and send the search request to the service platform 2. The player can output (such as display, present) content to the user. The content can include video, audio, comments, text data, and / or the like. The audience terminals (6A, 6B, …, 6N) can include an interface, and the interface can include an input element (touch screen). For example, the input element can be configured to receive user instructions, and the user instructions can cause the audience terminals (6A, 6B, …, 6N) to perform various operations, such as sending bullet comments, inputting comments, giving gifts, etc.
[0036] The host terminals (4A, 4B, …, 4M), the audience terminals (6A, 6B, …, 6N) and the service platform 2 can be connected through a network. The network can include various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, and / or proxy devices, etc. The network can include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links, and their combinations and / or the like. The network can include wireless links, such as cellular links, satellite links, Wi-Fi links, and / or the like.
[0037] It should be noted that Figure 1 the numbers of the host terminals and the audience terminals in are only illustrative and are not used to limit the patent protection scope of the present application. According to the actual situation, there can be any number of host terminals and audience terminals, and the host terminal and the audience terminal can be the same terminal.
[0038] Next, taking the service platform (server) 2 as the execution entity, the technical solutions of the present application will be introduced through multiple embodiments. It should be noted that these embodiments can be implemented in various different forms and should not be construed as being limited only to the embodiments described herein.
[0039] Embodiment 1 Figure 2The flowchart of the search method according to Embodiment 1 of the present application is schematically shown.
[0040] As Figure 2 shown, the search method may include steps S200 to S206, where: Step S200: Receive a search request sent by a client, where the search request carries an identifier of a target object.
[0041] Step S202: Based on the identifier of the target object, determine a target index entry from multiple index entries of an object inverted index. Among them, the object inverted index is constructed based on video contents of multiple live rooms, and the video contents correspond to one or more objects. Each index entry corresponds to an object and indicates an identifier of a live room associated with the object.
[0042] Step S204: According to the target index entry, determine an identifier of a target live room associated with the target object.
[0043] Step S206: Generate a target live room list according to the identifier of the target live room, and push the target live room list to the client.
[0044] In the search method provided in this embodiment, the server obtains video contents of multiple live rooms, and the video contents correspond to one or more objects. An object inverted index is constructed based on the relationship between the object and the live room. Among them, the object inverted index includes multiple index entries, each index entry corresponds to an object and indicates an identifier of a live room associated with the object. The server receives a search request from the client, and the search request carries an identifier of a target object. Based on the identifier of the target object, a target index entry is determined from multiple index entries of the object inverted index. According to the target index entry, an identifier of a target live room associated with the target object is determined. A target live room list is generated based on the identifier of the target live room and is recommended to the client as a search result. It can be seen that through understanding the video contents of the live room in the embodiment of the present application, the object being displayed or explained can be accurately identified, and an object inverted index is established based on the relationship between the object and the live room. When the user searches for a target object, the associated target live room can be quickly obtained, and a target live room list is generated as a search result and pushed to the user. That is, more intuitive, accurate and comprehensive search results can be provided, the problems of missed detection and misdetection can be alleviated, and thus the search experience and platform competitiveness can be effectively improved.
[0045] The following Figure 2 will elaborate in detail on each step in steps S200 to S206 and other optional steps.
[0046] Step S200 Receive a search request sent by a client, where the search request carries an identifier of a target object.
[0047] The object can be various types of content, such as: physical goods, virtual goods (such as game props, virtual skins, etc.), social accounts, media content (such as music, articles, pictures, etc.). Users can log in to and access the search interface of the content platform through a client (such as a host terminal, an audience terminal), and the search interface can be configured with a search box. Users can enter query text in the search box, and the query text can include the identifier of the target object. Among them, the target object can be an object that the user is interested in, and the identifier can be the name, ID, keyword, etc. of the object. Based on the query text, a search request can be generated and sent to the server. The server parses the search request and can obtain the identifier of the target object.
[0048] After determining the identifier of the target object, the server can directly match the identifier of the target object with the titles of all live rooms, find the live room with the highest keyword matching degree and return it to the client. However, due to situations such as the limited number of characters in the live room title and the inconsistency between the live room title and the content, it may not be possible to accurately identify the object corresponding to the live room only through title matching, which may result in the live rooms related to the target object not being retrieved (missed detection) or recommending live rooms unrelated to the target object (false detection), thus affecting the search experience. To improve the accuracy of search results, exemplary solutions will be provided below.
[0049] Step S202 , based on the identifier of the target object, determine a target index entry from multiple index entries of the object inverted index; wherein, the object inverted index is constructed based on the video content of multiple live rooms, the video content corresponds to one or more objects, and each index entry corresponds to an object and indicates the identifier of the live room associated with the object.
[0050] The object inverted index can be constructed by the server based on the video content of multiple live rooms. Exemplarily, the server can obtain the video content of all live rooms in real time, and the video content can include video (image frames) and audio (such as host voice, background music, etc.). Through technologies such as computer vision, machine learning, and deep learning, the video content is understood to accurately identify all the objects corresponding to the live rooms, such as Figure 3As shown in the figure. For example, live stream room A is displaying the product "air fryer". By extracting the image frames of this live stream room in real time and through image recognition, it can be determined that the product corresponding to live stream room A is "air fryer". Live stream room B is explaining the product "Bluetooth headset". By recognizing the voice of the host in this live stream room in real time, it can be determined that the product corresponding to live stream room B is "Bluetooth headset". Of course, the video content can correspond to one or more objects, that is, a live stream room can also display and / or explain multiple objects at the same time. For example, live stream room C is simultaneously displaying "Bluetooth headset" and "noise-canceling headset", and the host is explaining the differences between the two. After determining the objects corresponding to each live stream room, it can be determined in which live stream rooms each object appears specifically, and accordingly, an object inverted index can be established, as Figure 4 shown. The object inverted index can include multiple index entries, each index entry corresponding to an object and indicating the identifier of the live stream room associated with that object. As Figure 5 shown, the object inverted index can be stored through a hash table, which can improve the query efficiency. Subsequently, fast insertion and update can be achieved, and new entries can be processed quickly. In the hash table, each index entry is a key-value pair, where the key can be the identifier of the object and the value can be the identifier of the live stream room. Among them, the identifier of the live stream room can be ID, URL, host ID, etc.
[0051] Based on the identifier of the target object, the server can quickly determine the target index entry corresponding to the target object from multiple index entries of the object inverted index, and subsequently use it to index the associated target live stream room.
[0052] In this embodiment, constructing an object inverted index based on video content can efficiently, accurately, and comprehensively retrieve the live stream rooms associated with the target object, significantly improving the accuracy of search results.
[0053] Multiple exemplary solutions for constructing an object inverted index will be provided below.
[0054] In an alternative embodiment, as Figure 6 shown, the object inverted index can be constructed through the following steps: Step S600, obtain the video content of the multiple live stream rooms, where the video content includes image frames.
[0055] Step S602, based on the image frames of the multiple live stream rooms, determine the objects corresponding to the multiple live stream rooms.
[0056] Step S604, based on the objects corresponding to the multiple live stream rooms, establish the object inverted index.
[0057] Exemplarily, the server can obtain the video streams of multiple live rooms in real time, and use video decoding technology (such as FFmpeg) to decompose the video streams into continuous image frames. In some embodiments, the image frames can also be preprocessed, such as resizing, denoising, color enhancement, etc., to improve the accuracy of subsequent recognition. By analyzing the image frames, the objects corresponding to the live rooms can be determined. Exemplarily, the server can extract the text in the image frames through optical character recognition (OCR) technology and match the extracted text with a preset object library. The preset object library can include the identification, text introduction, etc. of each object. If the extracted text matches a certain object in the preset object library, it means that the object appears in the live room, and it can be determined as the object corresponding to the live room. In an alternative embodiment, the server can also process the image frames of the multiple live rooms through visual analysis technology or a pre-trained image recognition model to determine the objects corresponding to the multiple live rooms. Among them, the visual analysis technology can be: using feature extraction algorithms (such as SIFT, SURF, OR) to extract the key features in the image frames and match them with a preset object feature library. The image recognition model can be a convolutional neural network (CNN), a target detection model (such as YOLO, Faster R-CNN), etc. By inputting the image frames into a pre-trained image recognition model, the image recognition model can quickly and accurately identify the objects in the image frames after deeply understanding the image frames, and correspondingly output the identification or pictures of the objects. After determining the objects corresponding to each of the multiple live rooms, the server can use the identification of each object as a key and the identifications of all the live rooms associated with it as the corresponding value to establish an object inverted index in this way. Use a hash table or other efficient data structures (such as B-trees) to store the object inverted index to improve the query performance.
[0058] In this embodiment, through the in-depth understanding of the image frames, the objects in the live rooms can be efficiently and accurately identified, and a reliable and efficient object inverted index can be constructed, which can effectively improve the search accuracy and experience.
[0059] In an alternative embodiment, as Figure 7 shown, the object inverted index can be constructed through the following steps: Step S700, obtain the video content of the multiple live rooms, where the video content includes audio.
[0060] Step S702, based on the audio of the multiple live rooms, determine the objects corresponding to the multiple live rooms.
[0061] Step S704, based on the objects corresponding to the multiple live rooms, establish the object inverted index.
[0062] Exemplarily, the server can obtain the video streams of multiple live rooms in real time and extract the audio therefrom. The audio can include voice interaction, ambient sound effects, background music, etc. Among them, the voice interaction can be the conversation between the host, guests, and / or voice assistants, and may directly mention the object, for example: explaining the product. The ambient sound effects may be the sounds related to the object, such as the demonstration sound effects of the object, such as the tapping sound of a mechanical keyboard. The background music is the music played in the live room and may be related to the object, such as the advertising BGM. The server can use audio analysis technologies (such as speech recognition, sentiment analysis, etc.) combined with natural language processing technologies to process the audio, and can accurately identify the object corresponding to the live room. After determining the objects corresponding to each of the multiple live rooms, an object inverted index can be established based on the live rooms where each object has appeared, and the object inverted index can be stored through a hash table.
[0063] In this embodiment, through in-depth analysis of the audio, the object corresponding to the live room can be accurately determined, which is used to construct the object inverted index to improve the accuracy of subsequent searches.
[0064] In an alternative embodiment, as Figure 8 shown, the object inverted index can be constructed through the following steps: Step S800, obtain the video content of the multiple live rooms, where the video content includes image frames and audio.
[0065] Step S802, through image recognition of the image frames of the multiple live rooms, determine the first objects corresponding to the multiple live rooms.
[0066] Step S804, through audio analysis of the audio of the multiple live rooms, determine the second objects corresponding to the multiple live rooms.
[0067] Step S806, according to the first objects and the second objects corresponding to the multiple live rooms, determine the objects corresponding to the multiple live rooms.
[0068] Step S808, based on the objects corresponding to the multiple live rooms, establish the object inverted index.
[0069] Exemplarily, the server can obtain the video content of multiple live rooms in real time, extract image frames and audio therefrom. By processing the image frames through visual analysis techniques or pre-trained image recognition models, the first objects corresponding to the multiple live rooms can be determined. By processing the corresponding audio through audio analysis techniques combined with natural language processing techniques, the second objects corresponding to the multiple live rooms can be determined. In a live broadcast scenario, the host often shows and explains objects simultaneously to improve the live broadcast effect. If the display and explanation of the object are not synchronized, it may be due to network latency, the host not updating the displayed object in time, or the display of the object being earlier than the host's explanation. Of course, it may also be that the host only briefly mentions the object, and this object is not the main content of the live broadcast and should not be determined as the object corresponding to the live room. Therefore, the objects corresponding to the multiple live rooms can be determined according to the matching situation of the first objects and the second objects. After determining the objects corresponding to the multiple live rooms, an object inverted index can be constructed for subsequent searches.
[0070] In this embodiment, by combining image frames and audio, the accuracy of object recognition is further improved, a more reliable inverted index is established, and the search accuracy, efficiency, and experience are optimized.
[0071] In an alternative embodiment, as Figure 9 shown, step S806 may include: Step S900, for each live room, determine whether the corresponding first object and the second object match.
[0072] Step S902, in the case of a match, determine the first object or the second object as the object corresponding to the live room.
[0073] Step S904, in the case of a non-match, determine the confidence levels corresponding to the image frame and the audio respectively.
[0074] Step S906, in the case where the confidence level of the image frame is higher than the confidence level of the audio, determine the first object as the object corresponding to the live room.
[0075] Step S908, in the case where the confidence level of the audio is higher than the confidence level of the image frame, determine the second object as the object corresponding to the live room.
[0076] Exemplarily, for each live streaming room, it can be determined whether the first object and the second object match. If the first object and the second object match, it indicates that both information sources (image frame and audio) point to the same object. The server can directly determine the object corresponding to the live streaming room as the first object or the second object. In the case where the first object and the second object do not match, the server will calculate the confidence levels corresponding to the image frame and the audio respectively. The confidence level represents the reliability of the analysis result of the image frame or the audio. Among them, the confidence value of the image frame can be determined according to the resolution and clarity of the image frame, the position of the object in the image frame, the performance of the image recognition model, the performance of the visual analysis algorithm, etc. The confidence level of the audio can be calculated according to the audio signal-to-noise ratio, speech clarity and pronunciation quality, context information of the speech, emotion analysis result, mention frequency of the object, etc. If the confidence level of the image is higher than that of the audio, it indicates that the information provided by the image frame is more credible, and the server can choose to determine the first object as the object corresponding to the live streaming room. Conversely, if the confidence level of the audio is higher than that of the image, it indicates that the information provided by the audio is more reliable, and the server can choose to determine the second object as the object corresponding to the live streaming room.
[0077] In this embodiment, in the case where the analysis results of the image frame and the audio do not completely match, selecting a more reliable object for the live streaming room according to the confidence level can improve the intelligent decision-making ability.
[0078] The above-mentioned multiple embodiments introduce building an object inverted index by obtaining the video content of multiple live streaming rooms in real time. In practical applications, as the video content changes in real time, the corresponding objects will also change. Correspondingly, the object inverted index also needs to be dynamically updated to ensure data consistency and adapt to real-time user search requirements.
[0079] In an alternative embodiment, as Figure 10 shown, the object inverted index can be updated through the following steps: Step S1000, obtain the latest video content of the multiple live streaming rooms, where the latest video content includes the current image frame and / or the current audio.
[0080] Step S1002, based on the current image frame and / or the current audio of the multiple live streaming rooms, determine the objects currently corresponding to the multiple live streaming rooms.
[0081] Step S1004, obtain the object list of each live streaming room, where the object list includes the objects corresponding to the history of the live streaming room, and the object list is determined based on the historical image frame and / or the historical audio of the live streaming room.
[0082] Step S1006, based on the object list of each live streaming room, determine whether there are objects that appear for the first time among the objects currently corresponding to each live streaming room.
[0083] Step S1008, when it is determined that an object appears for the first time, update the object inverted index according to the object that appears for the first time and the corresponding live room increment, and add the object that appears for the first time to the object list of the corresponding live room.
[0084] Exemplarily, the server can obtain the latest video content of multiple live rooms in real time, and can extract the current image frame and / or current audio from the latest video content. According to the current image frame and / or current audio, combined with technologies such as computer vision, machine learning, deep learning, audio analysis, and natural language processing, the objects corresponding to the multiple live rooms currently can be determined. Obtain the object list of each live room, and the object list can include the objects corresponding to the history of the live room. The object list can be determined according to the historical image frame and / or historical audio of the live room. Based on the historical object list and the currently corresponding objects of each live room, it can be determined whether there is an object that appears for the first time in each live room. If not, there is no need to update the object inverted index. If there is an object that appears for the first time, the object inverted index can be incrementally updated according to the object that appears for the first time and the corresponding live room, and the object that appears for the first time can be added to the object list for subsequent judgment.
[0085] In this embodiment, by maintaining the object list, new objects that appear in the live room can be captured in real time, so as to dynamically update the object inverted index, ensure the timeliness and accuracy of the search results, and further optimize the search accuracy.
[0086] Step S204 , according to the target index entry, determine the identifier of the target live room associated with the target object.
[0087] Exemplarily, the value in the target index entry can be read to obtain the identifier of the target live room associated with the target object. As Figure 5 shown, if the identifier of the target object is headphones, it can be queried in the hash table that "headphones" = [2, 3], then the identifiers of the target live rooms are 2 and 3.
[0088] Step S206 , generate a target live room list according to the identifier of the target live room, and push the target live room list to the client.
[0089] Exemplarily, relevant information of the target live streaming room, such as the cover image of the live streaming room and the title of the live streaming room, can be obtained according to the identifier of the target live streaming room. Based on this information, a list of target live streaming rooms can be generated. The server can push the list of target live streaming rooms to the client. The client can display the list of target live streaming rooms to the user. In this way, after the user enters the query text, the target live streaming rooms corresponding to the query text can be directly seen. Compared with only returning the target object, the embodiments of the present application can greatly improve the search experience.
[0090] To make the present application easier to understand, the following provides Figure 11 an exemplary application.
[0091] Taking an e-commerce platform as an example in this exemplary application, the object can be a commodity, and the specific search method is as follows: S1: The user logs in to the e-commerce platform through the client and accesses the search interface, which is configured with a search box. The user enters the keyword of the commodity of interest (such as headphones) in the search box. The client generates a search request based on the user input and sends it to the server.
[0092] Among them, the server is used to: extract image frames of multiple live streaming rooms in real time, and through image recognition technology, identify the commodities in the image frames, as Figure 3 shown. Establish a commodity inverted index for the commodities appearing in the live streaming room, as shown in 4. Store the commodity inverted index through a hash table, as Figure 5 shown.
[0093] S2: The server parses the search request and looks up in the hash table based on the keyword of the commodity. If "headphones" = [2, 3] (target index entry), then the identifiers of the target live streaming rooms are 2 and 3. Return these two live streaming rooms to the client as search results.
[0094] In this exemplary application: By understanding the video content of the live streaming room, the commodities explained by the anchor are identified. Then when the user searches for the commodity name, the relevant live streaming rooms can be recommended to the user, which can improve the user search experience, improve the accuracy of the search results, and at the same time enhance the competitiveness of the e-commerce platform.
[0095] Embodiment 2 Figure 12 The block diagram of the search device according to Embodiment 2 of the present application is schematically shown. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and are executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 12As shown in the figure, the device 1000 may include: a receiving module 1100, a first determination module 1200, a second determination module 1300, and a pushing module 1400, where: The receiving module 1100 is configured to receive a search request sent by a client, and the search request carries an identifier of a target object; The first determination module 1200 is configured to determine a target index entry from multiple index entries of an object inverted index based on the identifier of the target object; wherein, the object inverted index is constructed based on video contents of multiple live rooms, the video contents correspond to one or more objects, and each index entry corresponds to an object and indicates an identifier of a live room associated with the object; The second determination module 1300 is configured to determine an identifier of a target live room associated with the target object according to the target index entry; The pushing module 1400 is configured to generate a target live room list according to the identifier of the target live room, and push the target live room list to the client.
[0096] As an optional embodiment, the object inverted index is constructed through the following operations: Obtain the video contents of the multiple live rooms, where the video contents include image frames; Determine the objects corresponding to the multiple live rooms based on the image frames of the multiple live rooms; Establish the object inverted index based on the objects corresponding to the multiple live rooms.
[0097] As an optional embodiment, determining the objects corresponding to the multiple live rooms based on the image frames of the multiple live rooms includes: Determine the objects corresponding to the multiple live rooms based on the image frames of the multiple live rooms through visual analysis technology or a pre-trained image recognition model.
[0098] As an optional embodiment, the object inverted index is constructed through the following operations: Obtain the video contents of the multiple live rooms, where the video contents include audio; Determine the objects corresponding to the multiple live rooms based on the audio of the multiple live rooms; Establish the object inverted index based on the objects corresponding to the multiple live rooms.
[0099] As an optional embodiment, the object inverted index is constructed through the following operations: Obtain the video contents of the multiple live rooms, where the video contents include image frames and audio; Determine a first object corresponding to the multiple live rooms by performing image recognition on the image frames of the multiple live rooms; Determine a second object corresponding to the multiple live rooms by performing audio analysis on the audio of the multiple live rooms; Determine the object corresponding to the multiple live rooms according to the first object and the second object corresponding to the multiple live rooms; Based on the objects corresponding to the multiple live rooms, establish an inverted index of the objects.
[0100] As an optional embodiment, determining the object corresponding to the multiple live rooms according to the first object and the second object corresponding to the multiple live rooms includes: For each live room, determine whether the corresponding first object and second object match; In the case of a match, determine the first object or the second object as the object corresponding to the live room; In the case of non - match, determine the confidence levels corresponding to the image frame and the audio respectively; In the case where the confidence level of the image frame is higher than that of the audio, determine the first object as the object corresponding to the live room; In the case where the confidence level of the audio is higher than that of the image frame, determine the second object as the object corresponding to the live room.
[0101] As an optional embodiment, the apparatus 1000 is further configured to further include: Obtain the latest video content of the multiple live rooms, where the latest video content includes a current image frame and / or current audio; Based on the current image frame and / or current audio of the multiple live rooms, determine the objects currently corresponding to the multiple live rooms; Obtain the object list of each live room, where the object list includes the objects corresponding to the history of the live room, and the object list is determined based on the historical image frames and / or historical audio of the live room; Based on the object list of each live room, determine whether there is an object that appears for the first time among the objects currently corresponding to each live room; In the case where it is determined that there is an object that appears for the first time, incrementally update the inverted index of the objects according to the object that appears for the first time and the corresponding live room, and add the object that appears for the first time to the object list of the corresponding live room.
[0102] Embodiment III Figure 13Schematically shown is a hardware architecture diagram of a computer device 10000 suitable for implementing a search method according to Embodiment 3 of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including a stand-alone server or a server cluster composed of multiple servers), etc. As Figure 13 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can be communicatively linked to each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the search method, etc. In addition, the memory 10010 may also be used to temporarily store various types of data that have been output or will be output.
[0103] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.
[0104] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.
[0105] It should be noted that Figure 13 only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0106] In this embodiment, the search method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of this application.
[0107] Embodiment Four The embodiments of this application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the search method in the embodiments are implemented.
[0108] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the search method in the embodiment. In addition, the computer-readable storage medium may also be used to temporarily store various data that have been output or are to be output.
[0109] Embodiment Five The embodiment of the present application further provides a computer program product, including a computer program, which when executed by a processor implements the method in the above embodiment.
[0110] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device, so that they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0111] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A search method, characterized in that, For a server, the method includes: Receiving a search request sent by a client, where the search request carries an identifier of a target object; Based on the identifier of the target object, determining a target index entry from multiple index entries of an object inverted index; wherein, the object inverted index is constructed based on video content of multiple live rooms, the video content corresponds to one or more objects, and each index entry corresponds to an object and indicates an identifier of a live room associated with the object; According to the target index entry, determining an identifier of a target live room associated with the target object; Generating a target live room list according to the identifier of the target live room and pushing the target live room list to the client.
2. The method according to claim 1, wherein The object inverted index is constructed through the following operations: Obtaining the video content of the multiple live rooms, where the video content includes image frames; Based on the image frames of the multiple live rooms, determining the objects corresponding to the multiple live rooms; Based on the objects corresponding to the multiple live rooms, establishing the object inverted index.
3. The method according to claim 2, wherein Based on the image frames of the multiple live rooms, determining the objects corresponding to the multiple live rooms includes: Based on the image frames of the multiple live rooms, determining the objects corresponding to the multiple live rooms through visual analysis technology or a pre-trained image recognition model.
4. The method according to claim 1, characterized in that The object inverted index is constructed through the following operations: Obtaining the video content of the multiple live rooms, where the video content includes audio; Based on the audio of the multiple live rooms, determining the objects corresponding to the multiple live rooms; Based on the objects corresponding to the multiple live rooms, establishing the object inverted index.
5. The method according to claim 1, wherein The object inverted index is constructed through the following operations: Obtaining the video content of the multiple live rooms, where the video content includes image frames and audio; Through image recognition of the image frames of the multiple live rooms, determining a first object corresponding to the multiple live rooms; Through audio analysis of the audio of the multiple live rooms, determining a second object corresponding to the multiple live rooms; According to the first object and the second object corresponding to the multiple live rooms, determining the objects corresponding to the multiple live rooms; Based on the objects corresponding to the multiple live rooms, establishing the object inverted index.
6. The method according to claim 5, characterized in that, According to the first object and the second object corresponding to the multiple live rooms, determining the objects corresponding to the multiple live rooms includes: For each live room, determining whether the corresponding first object and second object match; In the case of a match, determining the first object or the second object as the object corresponding to the live room; In the case of no match, determining the confidence levels corresponding to the image frames and the audio respectively; In the case where the confidence level of the image frames is higher than the confidence level of the audio, determining the first object as the object corresponding to the live room; In the case where the confidence level of the audio is higher than the confidence level of the image frames, determining the second object as the object corresponding to the live room.
7. The method according to claim 5, characterized in that The object inverted index is constructed through the following operations: Obtaining the latest video content of the multiple live rooms, where the latest video content includes current image frames and / or current audio; Determine the objects currently corresponding to the multiple live rooms based on the current image frames and / or current audio of the multiple live rooms; Obtain the object list of each live room, where the object list includes the objects historically corresponding to the live room, and the object list is determined based on the historical image frames and / or historical audio of the live room; Based on the object list of each live room, determine whether there are objects that appear for the first time among the objects currently corresponding to each live room; In the case where it is determined that there are objects that appear for the first time, incrementally update the object inverted index according to the objects that appear for the first time and the corresponding live rooms, and add the objects that appear for the first time to the object list of the corresponding live rooms.
8. A search device, characterized in that, For a server, the device includes: A receiving module, configured to receive a search request sent by a client, where the search request carries an identifier of a target object; A first determination module, configured to determine a target index entry from multiple index entries of an object inverted index based on the identifier of the target object; wherein, the object inverted index is constructed based on the video content of multiple live rooms, the video content corresponds to one or more objects, and each index entry corresponds to an object and indicates the identifier of the live room associated with the object; A second determination module, configured to determine the identifier of the target live room associated with the target object according to the target index entry; A pushing module, configured to generate a target live room list according to the identifier of the target live room and push the target live room list to the client.
9. A computer device, characterized in that, Includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein: The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to claims 1 to 7 are implemented.
Citation Information
Cited By
Multi-channel advertisement putting effect analysis method and system based on real-time calculation
CN121094888A