Video resource search method and related device

By combining a large text model and a PGC resource library, the system identifies users' natural language intent, enabling precise search and segment recommendation of video resources. This solves the problem of inaccurate searching in traditional methods and improves the user experience.

WO2026098409A1PCT designated stage Publication Date: 2026-05-15HUAWEI TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-11-03
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Traditional video resource search methods cannot support users in accurately searching for specific segments using natural language, nor can they meet users' fragmented viewing needs.

Method used

By combining textual and multimodal large models with internet and professionally generated content (PGC) resource libraries, it identifies users' natural language intent, accurately searches for and recommends video resource clips, and supports multi-turn dialogue and resource selection.

Benefits of technology

It enables users to search for specific video resource segments using any natural language, improving search accuracy and response speed, and meeting users' precise search needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025132209_15052026_PF_FP_ABST
    Figure CN2025132209_15052026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video resource search method and a related device. By utilizing rich UGC on the Internet, high-quality PGC in a video resource library, and the semantic understanding capability and content understanding capability of a large text model, users are allowed to accurately find video resources they want by means of conversational queries. In addition, after a video resource segment conforming to the intention of a user is determined by using Internet search results and PGC search results, a multi-modal large model can be used to further locate a finer-grained segment conforming to the intention of the user from the video resource segment, thereby meeting fragmented viewing needs of users.
Need to check novelty before this filing date? Find Prior Art

Description

Video resource search methods and related equipment

[0001] This application claims priority to Chinese Patent Application No. 202411600681.5, filed on November 8, 2024, entitled "Resource Search Method and Related Equipment", and to Chinese Patent Application No. 202411873273.7, filed on December 17, 2024, entitled "Video Resource Search Method and Related Equipment", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of electronic technology, and in particular to video resource search methods and related equipment. Background Technology

[0003] With the increasing popularity of voice-interactive products such as in-vehicle systems, smart screens, and smart speakers, users are increasingly inclined to use voice commands to search for and control video resources. Moreover, due to the rise of short videos and the increasing pressure of modern life, more and more users are accustomed to watching videos during fragmented time, leading to a growing demand for precise searches of video clips.

[0004] Traditional video resource searches rely more on explicit tags such as movie titles, actor names, narrator names, and movie genres. These tags require users to input them to find relevant video resources. They do not support users using more colloquial natural language for searching and cannot accurately find the specific segment a user wants within the video resources. Summary of the Invention

[0005] This application provides a video resource search method and related equipment, which can support users to search for video resources, audio resources and other resources using any natural language, and support users to accurately search for specific resource segments, thus meeting users' needs for precise resource search.

[0006] In a first aspect, embodiments of this application provide a video resource search method. This method can be applied to a terminal device and specifically includes the following steps: The terminal device receives first dialogue content input by a first user, the first dialogue content expressing a resource search intent in natural language. Based on the first dialogue content, the terminal device replies with a first resource fragment that matches the resource search intent. The first resource fragment is a video resource fragment, which is one or more episodes of a film or television work, or one or more film or television clips from an episode of a film or television work. Replying with the first resource fragment that matches the resource search intent includes displaying a first reply text and / or displaying a first service card. The first reply text describes the first resource fragment, and the first service card is used by the user to view the first resource fragment.

[0007] In the first aspect, the user can input the first dialogue content through a smart voice assistant, such as by speaking the first dialogue content aloud. However, the user can also input the first dialogue content via text input; this application embodiment does not limit the input method for the first dialogue content.

[0008] Implementing the method provided in the first aspect allows users to search for video resources using any natural language and to precisely search for specific resource segments, thus meeting users' needs for accurate resource searches.

[0009] In conjunction with the first aspect, in some embodiments, the rendered content on the first service card can also serve to describe the first resource fragment, and its rendered content may include one or more of the following: the name, type, release time, creator, cover image, etc. of the first resource fragment or its associated video resource.

[0010] In conjunction with the first aspect, in some embodiments, when the terminal device detects a user operation applied to the first service card, the terminal device can jump to the playback interface of the first resource segment and play the first resource segment. Besides clicking the first service card as user input, the terminal device can also receive user input to trigger the playback of the first resource segment through a continuation dialogue, and then jump to the playback interface of the first resource segment to play it. The continuation dialogue may include: first asking whether to play the first resource segment; then receiving the user's response to the inquiry. This continuation dialogue can be a voice dialogue or a text dialogue; this application embodiment does not limit the dialogue format. The implementation details of triggering the playback of the first resource segment can be found in the relevant content of the foregoing embodiments, and will not be repeated here.

[0011] In combination with the first aspect, in some embodiments, the terminal device may also send the first conversation content to the server and receive the indication information of the first resource segment from the server. The indication information of the first resource segment includes the identification information of the resource to which the first resource segment belongs and the positioning information of the first resource segment in its所属 resource. Among them, the positioning information includes the start time of the first resource segment. The identification information of the resource may be, for example, the resource name, and the positioning information of the first resource segment in its所属 resource may be the resource segment number, time range, etc. For example, the indication information of the first resource segment is composed of the resource name "Empresses in the Palace" and the resource segment number "Episode 63". Another example is that the indication information of the first resource segment is composed of the resource name "Infernal Affairs" and the time range "30:00-40:00". When using the time range as the positioning information of the first resource segment in its所属 resource, the start time of the first resource segment can be used as the key parameter indicating this time range, and the end time of the first resource segment can be defaulted to the end time of the resource to which the first resource segment belongs if not explicitly indicated.

[0012] In combination with the first aspect, in some embodiments, the terminal device may also receive the first reply text from the server.

[0013] In combination with the first aspect, in some embodiments, the terminal device may also receive the first picture from the server, and the first picture is generated according to the first resource segment. Specifically, the first picture may come from the first resource segment, such as the first frame picture of the first resource segment. In this way, the terminal device can use the first picture as the cover picture of the first service card when rendering and generating the first service card, improving the user attraction of the first service card.

[0014] In combination with the first aspect, in some embodiments, the terminal device may also receive the first conversation content input by the second user, and according to the first conversation content input by the second user, the terminal device replies to the second user with the second resource segment that meets the resource search intention. The second resource segment is one or more episodes in a film or television work, or one or more film or television segments in an episode of a film or television work. The second user is different from the first user, and the second resource segment is different from the first resource segment. Among them, the second resource segment may also be one or more episodes in a film or television work, or one or more film or television segments in an episode of a film or television work. Among them, replying to the second user with the second resource segment that meets the resource search intention may include one or more of the following methods: displaying the second reply text, displaying the second service card. The second reply text can be used to describe the second resource segment and may include the recommended reasons, plot introductions, etc. of the second resource segment. The second service card can be used for the user to watch the second resource segment.

[0015] In conjunction with the first aspect, in some embodiments, the memory information or personal profiles of the first user and the second user on the terminal device may differ. Memory information may include the user's video browsing history, video collection history, and other usage traces that can indicate the user's preferences for film and television resources. The personal profile may be a user profile determined by the cloud side based on user data uploaded by the terminal (such as gender, age, occupation, etc.) and / or the aforementioned memory information, used to indicate the user's film and television preferences. User film and television preferences can be broad, including preferences for film and television content, preferences for film and television creators (such as actors), preferences for film and television sources, preferences for playback resolution, etc.

[0016] In conjunction with the first aspect, in some embodiments, the second resource segment may differ from the first resource segment in, but is not limited to, the following ways: the second resource segment has different video content from the first resource segment, or the second resource segment has the same content as the first resource segment but belongs to different resources, or the second resource segment has the same content as the first resource segment but has different clarity.

[0017] In conjunction with the first aspect, in some embodiments, there may be multiple first resource fragments that match the resource search intent. In this case, after replying to the first user with the first resource fragments that match the resource search intent, the terminal device may further receive second dialogue content input by the first user. The second dialogue content expresses, in natural language, the resource selection intent to select a third resource fragment from the multiple first resource fragments. Then, the terminal device can select the third resource fragment from the multiple first resource fragments according to the second dialogue content and jump to the playback interface of the third resource fragment to play the third resource fragment.

[0018] In conjunction with the first aspect, in some embodiments, the first resource fragment conforming to the resource search intent can be determined by the terminal device based on the content of the first dialogue, especially when the terminal device has strong computing power. Alternatively, the first resource fragment conforming to the resource search intent can also be determined by the cloud-side server based on the content of the first dialogue and fed back to the terminal device, which then replies with the first resource fragment to the user.

[0019] Second aspect, an embodiment of the present application provides a video resource search method, which can be applied to a server. Specifically, it may include the following steps: The server receives the first conversation content from the terminal device. The first conversation content expresses the resource search intention in natural language. The server determines a first resource segment that meets the resource search intention for the first conversation content and returns the indication information of the first resource segment to the terminal device. Among them, the first resource segment is a video resource segment, and the first resource segment is one or more episodes of a film and television work, or one or more film and television segments in an episode of a film and television work; the indication information of the first resource segment includes the identification information of the resource to which the first resource segment belongs, and the positioning information of the first resource segment in its affiliated resource. Among them, the positioning information includes the start time of the first resource segment.

[0020] In the second aspect, the indication information of the first resource segment includes the identification information of the resource to which the first resource segment belongs, and the positioning information of the first resource segment in its affiliated resource. Among them, the positioning information includes the start time of the first resource segment. The identification information of the resource can be, for example, the resource name. The positioning information of the first resource segment in its affiliated resource can be the resource segment number, time range, etc. For example, the indication information of the first resource segment is composed of the resource name "Empresses in the Palace" and the resource segment number "Episode 63". Another example is that the indication information of the first resource segment is composed of the resource name "Infernal Affairs" and the time range "30:00 - 40:00". When using the time range as the positioning information of the first resource segment in its affiliated resource, the start time of the first resource segment can be used as the key parameter indicating this time range. Without clear indication, the end time of the first resource segment can be defaulted to the end time of the resource to which the first resource segment belongs.

[0021] Implementing the method provided in the second aspect, the inference operation steps completed based on the text large model and the multimodal large model can be executed by a cloud - side server with stronger computing power, which can reduce the requirements for the computing power of the terminal device on the end - side, reduce the end - side power consumption, etc., and reduce the user waiting time.

[0022] Combined with the second aspect, in some embodiments, the server can also return a first reply text to the terminal device. The first reply text is used to describe the first resource segment.

[0023] Combined with the second aspect, in some embodiments, the server can also return a first picture to the terminal device. The first picture can be generated according to the first resource segment. Specifically, the first picture can come from the first resource segment, such as the first - frame picture of the first resource segment. In this way, the terminal device can use the first picture as the cover picture of the first service card when rendering and generating the first service card, improving the user attraction of the first service card.

[0024] In conjunction with the second aspect, in some embodiments, the server determines a first resource fragment that matches the resource search intent based on the first dialogue content. Specifically, this may include: the server using PGC search results and Internet search results associated with the first dialogue content to determine a first resource fragment that matches the resource search intent based on the first dialogue content.

[0025] Specifically, the internet search results associated with the first dialogue content are obtained by searching the internet based on the first dialogue content. The professionally produced content (PGC) search results associated with the first dialogue content are obtained by searching the PGC resource library based on the first resource identifier. The first resource identifier is extracted from the internet search results, and the PGC resource library records the resource fragments included in different resources. In the specific implementation, the server can use the first customized prompt guiding text model to extract resource identifiers such as video resource names from the internet search results.

[0026] In conjunction with the second aspect, in some embodiments, compared to the resource fragments recorded in the PGC resource library, the first resource fragment can be a shorter resource fragment, that is, a more granular resource fragment. To this end, the server can first determine a third resource fragment that matches the resource search intent based on internet search results and PGC search results. The duration of the third resource fragment can be consistent with the duration of the resource fragments recorded in the PGC resource library. Then, the server can further process the third resource fragment, internet search results, and PGC search results using a multimodal large model to further locate the first resource fragment from the third resource fragment. In this way, the first resource fragment can meet the user's fragmented viewing needs.

[0027] In conjunction with the second aspect, in some embodiments, the server determines a first resource fragment that matches the resource search intent based on the first dialogue content. Specifically, this may include: the server analyzing internet search results and PGC search results based on the first user's user profile to determine the first resource fragment. In this way, different users can receive different recommended resource fragments, even if different users input the same dialogue content.

[0028] In conjunction with the second aspect, in some embodiments, before the server determines a first resource fragment that matches the resource search intent based on the first dialogue content, the server also uses PGC search results to filter out inaccurate content in Internet search results to improve the accuracy of resource recommendations and responses.

[0029] In conjunction with the second aspect, in some embodiments, there may be multiple first resource fragments that match the resource search intent. After the server returns indication information for the first resource fragment to the terminal device, the server may also receive second dialogue content from the terminal device, which expresses the resource selection intent in natural language, and determine a third resource fragment that matches the resource selection intent from the multiple first resource fragments based on the second dialogue content, and then return indication information for the third resource fragment to the terminal device. The indication information for the third resource fragment is used to indicate the third resource fragment among the multiple first resource fragments.

[0030] Thirdly, this application provides a terminal device, which includes one or more processors and one or more memories; wherein the memories are coupled to the processors, and the one or more memories are used to store computer programs. When the processor executes the computer programs, it can implement the steps performed by the terminal device in the video resource search method described in the first aspect or any possible implementation of the first aspect.

[0031] Fourthly, this application provides a server comprising one or more processors and one or more memories; wherein the memories are coupled to the processors, and the one or more memories are used to store a computer program that, when executed by the processor, can implement the steps performed by the server in the video resource search method described in the second aspect or any possible implementation of the second aspect.

[0032] Fifthly, embodiments of this application provide a chip system applied to an electronic device. The chip system includes one or more processors, which are used to invoke computer instructions to implement the steps performed by the terminal device in the video resource search method described in the first aspect or any possible implementation of the first aspect.

[0033] In a sixth aspect, embodiments of this application provide a chip system applied to an electronic device. The chip system includes one or more processors, which are used to invoke computer instructions to implement the steps performed by the server in the video resource search method described in the second aspect or any possible implementation of the second aspect.

[0034] In a seventh aspect, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the steps performed by the terminal device in the video resource search method described in the first aspect or any possible implementation of the first aspect.

[0035] Eighthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the steps performed by the server in the video resource search method described in the second aspect or any possible implementation of the second aspect.

[0036] Ninthly, this application provides a computer program product containing instructions, which, when executed by a processor, can implement the steps performed by the terminal device in the video resource search method described in the first aspect or any possible implementation of the first aspect.

[0037] In a tenth aspect, this application provides a computer program product containing instructions that, when executed by a processor, can implement the steps performed by the server in the video resource search method described in the second aspect or any possible implementation of the second aspect.

[0038] In the eleventh aspect, this application provides an end-to-cloud collaborative system, which may include the terminal device described in the third aspect and the server described in the fourth aspect. Attached Figure Description

[0039] Figures 1A and 1B exemplarily illustrate typical applications of embodiments of this application in voice dialogue scenarios;

[0040] Figure 2 illustrates an intelligent system 20 based on a large model for resource search, selection, and online playback provided in an embodiment of this application;

[0041] Figure 3 illustrates the overall flow of the resource search method provided in Embodiment 1;

[0042] Figure 4 shows an example of the resource search method provided in Embodiment 1;

[0043] Figure 5 shows the first custom prompt used to guide the large text model;

[0044] Figure 6 illustrates the overall flow of the resource search method provided in Embodiment 2;

[0045] Figure 7 shows an example of the resource search method provided in Embodiment 2;

[0046] Figure 8 shows the second custom prompt used to guide the large text model;

[0047] Figure 9 illustrates the overall flow of the resource search method provided in Embodiment 3;

[0048] Figure 10 shows an example of the resource search method provided in Embodiment 3;

[0049] Figure 11 shows the third custom prompt used to guide the large text model;

[0050] Figure 12 illustrates the overall flow of the resource search method provided in the embodiments of this application;

[0051] Figure 13 illustrates the process of implementing the resource search method provided in this application embodiment through edge-cloud collaboration;

[0052] Figure 14 illustrates the terminal device provided in an embodiment of this application;

[0053] Figure 15 shows a server provided in an embodiment of this application. Detailed Implementation

[0054] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be a limitation of this application.

[0055] Supporting users to search for video resources using natural language can improve the user experience.

[0056] Currently, a technology for searching video resources via voice interaction uses a large model to identify user intent in resource queries, generate movie / TV show titles as search terms, and return these search terms to a video resource database via a JSON structure. The final search results are then displayed and presented via voice. However, this technology relies entirely on a static knowledge base provided by the large model to retrieve video resource names as search results. Updates to this knowledge base are delayed due to training limitations, affecting the timeliness of search results and the accuracy of finding emerging and popular videos. Furthermore, this technology may lead to the large model misinterpreting certain plot details due to the lack of such details in the knowledge base, potentially resulting in fabricated video resources or storylines.

[0057] Another technique based on natural language video segment search uses a knowledge graph and a BERT model to obtain natural language vector representations of the query, and encodes specific videos online into vector representations of video segments. It then calculates the Gumbel-Softmax score of these two vectors to obtain the corresponding video segments and presents them to the user. However, this technique implicitly calculates the most similar video segment by encoding the question and specific video into vector representations of the same dimension. This makes it difficult to intuitively provide reasons for recommending video segments or explicit query logic, which can easily mislead users. Furthermore, since multiple vector calculations inevitably have a maximum value, and the maximum value does not necessarily represent similar segments, it may still provide a "most relevant" segment even when the user queries irrelevant information (i.e., no results), significantly impacting accuracy. Moreover, this technique relies on an offline knowledge graph database built on Wikipedia, and is affected by Wikipedia's timeliness and coverage, failing to meet users' broad resource query needs and achieving the same effect as internet resource searches. Additionally, this technique requires online frame-by-frame analysis, filtering, and selection of video resources, resulting in slow response times and high latency.

[0058] This application provides a resource search method that allows users to search for resources such as video and audio resources using any natural language, and also allows users to search for specific resource fragments, thus meeting users' needs for precise resource search.

[0059] Figures 1A and 1B exemplify a typical application of the embodiments of this application in a voice dialogue scenario. In this voice dialogue scenario, users can search for video resources they want to learn about or watch by conversing with a voice assistant (such as "Xiaoyi"), and continue the conversation to select and control the playback of video resources. Wherein:

[0060] (1) Example in Figure 1A:

[0061] The first step is for the user to open the voice assistant and speak their question. This question can be used to search for a single video resource and can be any natural language voice input, such as "Which episode features the blood-drop kinship test in Zhen Huan?" This question can be displayed in the voice assistant's dialogue interface 11. In response to the question, the voice assistant can display a reply text 15 in its dialogue interface 11, which describes which episode the user wants to search for, such as "The Legend of Zhen Huan is a TV series, not a movie, and the blood-drop kinship test scene occurs in episode 63..." The voice assistant can also read the reply text 15 aloud. This voice reading can be triggered automatically, such as displaying and reading the reply text 15 line by line; it can also be triggered by the user, such as by clicking the voice reading control in the dialogue interface 11.

[0062] Furthermore, in response to this issue, the voice assistant can display a service card 16 in the dialogue interface 11. Service card 16 is the service card for the retrieved target video resource. The service card is an interface display format that places important application information or operations at the forefront, achieving direct service access and reducing user experience layers. Clicking service card 16 will trigger playback of the target video resource. The target video resource is the video resource that the user intends to learn about and / or view based on the identified user intent, such as episode 63 of the TV series "Empresses in the Palace".

[0063] The second step is for users to continue their conversation with the voice assistant to control the terminal device to play the target video resource.

[0064] For example, the voice assistant can continue to ask the user whether to play the target video resource. This inquiry can be made via voice, and / or by displaying a question text after the response text 15, such as "Found a relevant video, would you like to play it?". To this inquiry, the user can give an affirmative reply, such as "Yes, please play," which will redirect to the playback interface 17 of the target video resource, triggering its automatic playback. The playback interface 17 can be the video playback interface of a video application that provides the target video resource.

[0065] (2) Example in Figure 1B:

[0066] The first step is for the user to open the voice assistant and speak their question. This question can be used for fuzzy searching or recommendations of resources. It can be any natural language voice input, such as "Recommend three movies suitable for family reunions." This question can be displayed in the voice assistant's dialogue interface 11. In response to the question, the voice assistant can display a reply text 18 in the dialogue interface 11. This reply text describes several video resources that match the user's intent, including three movies suitable for family reunions, such as "Inside Out: This is an animated film produced by Pixar, through…", "The Pursuit of Happiness: This is an inspirational film starring Will Smith, with…", and "Les infideles: This is a French film about a naughty boy and…". The voice assistant can also read the reply text 18 aloud.

[0067] In addition, in response to this question, the voice assistant can also display multiple service cards 19 in the dialogue interface 11. These service cards 19 are service cards for multiple target video resources that have been retrieved. These target video resources are multiple video resources that the user intends to learn about and / or play based on the question, such as "Inside Out," "The Pursuit of Happyness," and "The Mischievous Boys of Paris." The user can choose to click on a specific service card 19 to trigger the playback of the selected video resource.

[0068] The second step is for users to continue their conversation with the voice assistant to select and control video resources.

[0069] For example, the voice assistant can continue to ask the user which video resource they want to play. This question can be asked verbally, and / or by appending a question text after the response text 15, such as "Find these videos, just tell me which one you want to watch." The user can respond verbally to this question, such as saying "Play the one starring Smith," to complete the video resource selection and be redirected to the playback interface 21 of that video resource, triggering its automatic playback.

[0070] As can be seen from the examples in Figures 1A and 1B, the embodiments of this application support users to conduct multi-turn dialogues with a voice-based intelligent assistant based on any natural language, in order to search for the video resources or video resource segments they want, further select the video resources they want to watch, and control the playback of the video resources they want to watch.

[0071] Moreover, this application embodiment combines the intent recognition and semantic understanding capabilities of the text big model with the high-quality PGC from the rich resource library of user-generated content (UGC) and professionally-generated content (PGC) on the Internet. While expanding the query scope and improving the timeliness of content through UGC, it ensures the professionalism and accuracy of query results by combining PGC.

[0072] Figure 2 illustrates an intelligent system 20 based on a large model for resource searching, selection, and online playback, provided in an embodiment of this application. The resources in Figure 2 are video resources, for example.

[0073] As shown in Figure 2, the intelligent system 20 mainly includes three modules: an intent recognition module, a resource search module, and a resource selection module. Among them:

[0074] 1. Functions of the intent recognition module:

[0075] 1.1 Obtain user profiles from the user profiling platform based on user identifier (ID) and application scenario identifier (ID).

[0076] A user profiling platform can be a cloud server that stores and manages user profiles of many users across different application scenarios. A user profile in a specific application scenario reflects that user's preferences and behaviors within that scenario. User profiles help intelligent systems understand user intent and predict user needs.

[0077] 1.2 Obtain the context of the multi-turn dialogue between the user and the voice assistant, that is, obtain the content of the multi-turn dialogue.

[0078] 1.3 The intent recognition capability based on the large text model can identify the user intent expressed by the user's statements in the dialogue scenario. The user intent may include: resource search intent and resource selection intent.

[0079] If the user's intent is resource search, then the resource search module will perform the resource search; if the user's intent is resource selection, then the resource selection module will perform the resource selection. Resource selection generally follows the resource search, and the resource search provides the user with multiple resources that match the user's search intent. The user then needs to further select the resource to play from these multiple resources. In multi-turn dialogue scenarios, the resource search dialogue is usually the preceding dialogue to the resource selection dialogue.

[0080] 2. Functions of the resource search module:

[0081] 2.1 Based on the question, request the search engine to obtain the Internet search results returned by the search engine, such as "search result 1", "search result 2", ..., "search result n".

[0082] Search engines can search the internet for content to obtain search results related to a question. The internet contains a lot of user-generated content (UGC) and some professionally generated content (PGC). The former has broad coverage and good timeliness; the latter provides more professional and high-quality content.

[0083] 2.2 Combining search results from the Internet with customized prompts, the content understanding capabilities of the text big data model are used to extract the names of video resources related to the problem from the search results, such as "movie name 1", "movie name 2", ..., "movie name n".

[0084] A custom prompt is used to guide the large text model to understand search results from the internet and output video resource names. This custom prompt will be described in detail in subsequent examples; it will not be elaborated upon here.

[0085] 2.3 Using the video resource names obtained in 2.2, query the PGC resource library to obtain the PGC search results for the video resource entities represented by the video resource names (such as "Video Resource Entity 1", "Video Resource Entity 2", ..., "Video Resource Entity n") in the PGC resource library. The PGC search results for a video resource entity may include some descriptive information (or feature information) of the video resource entity, such as film and television work introduction, actors, film and television genre, episode plot, time-sharing plot, etc.

[0086] Figure 2 illustrates the content of video resource entities in the PGC resource library, such as "Film Introduction: This film tells the story of...", "Actor Information: Starring Will Smith", "Film Type: Inspirational / Drama", "Episode Synopsis: Episode 1 synopsis is... Episode 2 synopsis is...", "Time-Based Synopsis: Episode 1 01:00-10:00 tells the story of... Episode 2 10:00-20:00 tells the story of...". Time-based synopsis is a more granular segment of the resource story than episode synopsis, and can be a segment of video resources of several minutes or even several seconds within a single episode video resource.

[0087] 2.4 Based on a text-based large-scale model, the system first analyzes internet search results and PGC search results to filter out inaccurate and low-quality content. Then, it identifies video resources that match the user's intent based on user dialogue content. The filtered search results are then used to summarize the text, resulting in film and television introductions and recommendation reasons, as shown in the aforementioned reply texts 15 and 18. Simultaneously, based on a multimodal large-scale model, inference operations are performed on the filtered search results and video resources that match the user's intent to locate video resource segments that match the intent. When locating video resource segments that match the user's intent, the multimodal large-scale model can also consider user profiles, ensuring that the located video resource segments are also closely aligned with user preferences.

[0088] After identifying video resources and video resource segments that match the user's intent, a service card for the video resource can be rendered on the client side. When the user clicks on the service card, they can be redirected to the video resource playback interface to watch the automatically playing video resource segment.

[0089] 3. Function of the resource selection module: Combining the aforementioned search results, the context of the previous dialogue, and the position of the video resource's service card in the previous dialogue interface, the module uses the semantic understanding capability of the text big data model to understand which video resource the user has selected, and finally jumps to the playback interface of that video resource to automatically play the video resource.

[0090] In the intelligent system 20, the intent recognition module, resource search module, and resource selection module can be deployed on a cloud-side server. The powerful computing capabilities of the cloud-side server can better ensure their smooth operation and performance, and also shorten inference computation time, saving users waiting time. Of course, the intent recognition module, resource search module, and resource selection module can also be deployed on edge devices such as mobile phones, tablets, in-vehicle devices, and smart screens, especially when these edge devices have strong computing capabilities.

[0091] Based on the foregoing, the resource search method provided by this application is described in detail below through several embodiments. The first embodiment focuses on a solution for accurately supplying video resource segments in response to natural language questions in a dialogue scenario; the second embodiment focuses on a solution for targeted supply of video resource segments by combining user profiles; and the third embodiment focuses on a solution for supporting users to query resources, select resources, and control resource playback in multi-turn dialogue scenarios.

[0092] In this application's embodiments, the term "question" or "user question" refers to a query in the field of natural language processing. Its meaning is not limited to a question; broadly speaking, it can be a task description. In the following text, "question" or "user question" can be replaced with user dialogue or user dialogue content.

[0093] All steps in the resource search method described in the following embodiments can be executed by terminal devices such as mobile phones, in-vehicle devices, and smart screens, especially when these terminal devices have strong computing capabilities. This resource search method can also be executed through edge-cloud collaboration. For example, human-computer interaction steps such as receiving queries through a voice assistant, outputting reply text, displaying resource service cards, and playback navigation can be executed by terminal devices such as mobile phones, in-vehicle devices, and smart screens. Meanwhile, the inference operations based on text-based large models and multimodal large models can be executed by cloud-based servers with stronger computing power. This reduces the computing power requirements of the terminal devices, lowers power consumption, and reduces user waiting time. In the edge-cloud collaboration approach, communication is required between the terminal devices and the cloud servers. The terminal needs to transmit the data that the text-based large model and multimodal large model need to process to the cloud, and the cloud needs to return the inference operation results to the terminal, such as which video resources match the user's intent, which video resource the user intends to play, etc. The following embodiments illustrate the resource search method using an edge-cloud collaboration approach.

[0094] Example 1

[0095] Figure 3 illustrates the overall flow of the resource search method provided in Embodiment 1. The details are as follows.

[0096] S11. The terminal device can receive a question uttered by the user and transmit the question to the server.

[0097] This question is expressed in natural language and is used to query resources. The question may include a spoken description of the resource to be queried, which may describe one or more of the following information: the creator of the resource (including actors, directors, singers, composers, authors, etc.), creation time, release time, resource type, and resource content, etc.

[0098] Resource types may include: video resources, audio resources, text resources, image resources, etc. Video resources may include, for example, movies, TV series, variety shows, and other film and television works; audio resources may include, for example, audiobooks, operas, radio broadcasts, etc.; text resources may include, for example, e-novels, e-journals, etc.; image resources may include, for example, e-magazines, etc. These examples are merely for explaining embodiments of this application and should not be construed as limiting the scope of the application.

[0099] In this embodiment, resource content can refer to the data characteristics of a resource, such as voice features, image features, text features, etc., or it can refer to the meaning expressed by the resource, such as plot, paragraph summary, lyrics, etc. For video resources, resource content can refer to the plot of the video resource; for audio resources, resource content can refer to the audio features (such as song melody), lyrics, lines, etc.; for text resources, resource content refers to the text content; for image resources, resource content can refer to the image content, that is, the meaning expressed by the image. The resource content of other resources can be defined similarly, and will not be listed here.

[0100] After receiving the user's input question, the terminal device can perform text preprocessing, query segmentation, query error correction, query alignment, query expansion, and other processing on the question to obtain the processed question, and then transmit the processed question to the server for processing.

[0101] S12. Based on the question, the server identifies the user's intent as a resource search intent, and searches the Internet based on the question, obtaining multiple Internet search results.

[0102] Specifically, you can request internet search engines to obtain these multiple internet search results. Internet search results can be in webpage format and may include text results, image results, etc.

[0103] The internet is characterized by the timeliness and breadth of its content, and as a constantly evolving and open system, it is not static but continuously expanding in scale. This allows users to search for newer and more content. However, much of the content on the internet is user-generated content (UGC), the professionalism and accuracy of which are not guaranteed. To address this, this application's embodiments will incorporate a professionally generated content (PGC) resource library to ensure the accuracy of resource searches, as will be explained later.

[0104] S13. The server can use a large text model to extract one or more resource entities from multiple Internet search results.

[0105] In this embodiment of the application, the essence of the resource entity is information, which mainly refers to resource identifiers such as resource name, and may also include auxiliary information such as publication time and creator.

[0106] For video resources, the resource name can refer to the title of a film or television work (such as a movie or TV series), the release date can refer to the release date of the film or television work, and the creator can include the director, actors, etc. For audio resources, the resource name can refer to the title of an audio work (such as an audiobook), the release date can refer to the release date of the audio work, and the creator can include narrators, voice actors, etc. For image resources, the resource name can refer to the title of an image work (such as the title of a painting or pictorial), the release date can refer to the publication date of the image work, and the creator can include painters, etc. For text resources, the resource name can refer to the title of a text work (such as a novel), the release date can refer to the publication date of the text work, and the creator can include authors, etc.

[0107] In its implementation, the terminal device can execute S13 using a first customized prompt and a large text model. The first customized prompt can be used to guide the large text model to extract resource entities that match the user's intent from the aforementioned multiple internet search results. The implementation of the first customized prompt will be explained in detail below, and will not be elaborated on here.

[0108] S14. The server can use the resource entity obtained in S13 to search in the PGC resource library and obtain the PGC search results for that resource entity.

[0109] A PGC search result for a resource entity primarily includes: descriptions of each resource segment within that entity, such as episode summaries or time-based synopses for a film or television work, or chapter content for an audiobook. A PGC search result for a resource entity may also include: an introduction to the resource entity, the creator, the resource type, etc.

[0110] The PGC resource library provides more professional and high-quality content. Compared to UGC on the internet, PGC is more accurate and professional. Furthermore, the PGC resource library in this embodiment records descriptions of resource fragments (such as time-sharing plots of film and television works) for each different resource entity, supporting the text-based large model in locating resource fragments that match the user's intent. This avoids the large model experiencing "illusion" due to its inability to identify resource fragments. Here, "illusion" refers to the large model fabricating non-existent resource fragments, such as fabricating plot segments.

[0111] The PGC resource library can be a PGC resource library on a cloud-based PGC server. When a user queries resources, the resource search module can retrieve PGCs from the PGC resource library online, such as resource descriptions, descriptions of resource fragments, resource types, publication times, creators, etc.

[0112] S15. The server can utilize the content understanding capabilities of the large text model to filter the aforementioned internet search results and PGC search results, removing low-quality and inaccurate content, and determine the resource fragment A that matches the user's intent based on the filtered internet search results and PGC search results.

[0113] Correspondingly, the server can also transmit indication information for resource segment A to the terminal device, such as the name of the film / TV work, the number of episodes, and the time range of the video segment (e.g., "30:00-40:00"). In this way, the terminal device can render a resource service card based on this indication information to display the resource service card in the dialogue interface of the voice assistant. Furthermore, when the user triggers playback of resource segment A, the terminal device uses the indication information of resource segment A to send a request to the video application's server to request the acquisition of resource segment A and play resource segment A on the terminal side.

[0114] Furthermore, the server can also transmit an introductory image for resource fragment A to the terminal device. This allows the terminal device to use the introductory image as the cover image for the corresponding resource service card, thereby enhancing the card's appeal to users. The introductory image for resource fragment A can be extracted by the server from resource fragment A. For example, the introductory image for resource fragment A could be the first frame of resource fragment A, such as the first frame of the 30:00-40:00 segment in episode 9 of "Conquest".

[0115] In S15, internet search results can help cover more and newer content, while PGC search results can help ensure content accuracy.

[0116] S16. The server can utilize the semantic and content understanding capabilities of the large text model to comprehensively analyze the aforementioned internet search results and PGC search results, generating a response text tailored to the question. The server can then transmit this response text to the terminal device, allowing the terminal device to display the response text to the user within the voice assistant's dialogue interface. The terminal device can also read the response text aloud.

[0117] Here, the internet search results used to generate the response text can be either the internet search results obtained in S12 directly, or internet search results filtered by PGC search results. This filtering can remove inaccurate content from the internet search results, thereby improving the accuracy of the response.

[0118] This response text can be used to describe resource A or resource fragment A. Specifically, the response text may include a brief introduction to resource A or resource fragment A, reasons for recommendation, etc., so that users can better understand resource A, resource fragment A, and the reasons for the response.

[0119] For example, in response to the question "Which episode features the blood test in 'The Legend of Zhen Huan'?", the reply could be: "'The Legend of Zhen Huan' is a TV series, not a movie. The blood test scene occurs in episode 63. In this episode, Zhen Huan is accused by Consort Qi of having an improper relationship with Imperial Physician Wen, which leads to..."

[0120] When generating recommendation reasons and summary introductions, internet search results help to broaden the scope of content, while PGC search results help to improve the accuracy of content.

[0121] S17. Based on the instruction information of resource fragment A, the terminal device can also display a resource service card for resource fragment A in the dialogue interface of the voice assistant. The rendered content of the resource service card can be generated based on the instruction information of resource fragment A. The user can click on the resource service card to view the automatic playback of resource fragment A.

[0122] Furthermore, by combining the filtered internet search results, PGC search results, and resource fragment A (i.e., continuous image frames), the server can leverage the multimodal large model's ability to process multimodal information to parse resource fragment A. This allows for the further localization of resource fragments within fragment A that better match the user's intent, particularly identifying the start time of the resource fragment. Compared to the resource fragments recorded in the PGC resource library, this further localized fragment is a finer-grained fragment identified by the multimodal large model, containing little or no redundant fragments related to other irrelevant issues. This enables users to more precisely query smaller resource fragments, meeting their needs for fragmented resource acquisition.

[0123] Accordingly, the server can transmit instruction information for resource fragment B to the terminal device, enabling the terminal device to render and generate a resource service card for resource fragment B. Users can then click on this resource service card to view the autoplay of resource fragment B. Furthermore, the server can transmit the first frame image of resource fragment B to the terminal device, allowing the terminal device to use this first frame image as the cover image for the resource service card, thereby increasing the card's appeal to users.

[0124] Resource service cards for resource fragments A and B can also be displayed in other locations, such as the message center, the negative one screen, the homepage of certain applications, the desktop, etc., and the resource service cards can even be transformed into other user interface forms. This application embodiment does not limit this.

[0125] More refined resource clip editing can also occur offline to reduce user waiting time online. First, large models can be used to identify video resources, generating finer-grained (minute-level, even second-level) resource clips along with descriptions. This is a time-consuming process; completing it offline reduces resource query latency. Second, during offline resource clip editing, user history questions and user profiles can be combined to generate resource clips that users might be interested in. These clips can then be pushed to users in the message center, the negative one screen, and the application homepage, increasing user engagement and improving user experience.

[0126] The resource search method provided in Embodiment 1 will be further explained in detail below using video resources as an example, with reference to Figure 4.

[0127] In the example in Figure 4, the question is "In which episode does Liu Huaqiang buy watermelons?" A solution for providing precise video resource clips to address this question could include the following steps:

[0128] Steps 1 to 4 primarily introduce the process of resource searching based on the intent recognition and content understanding capabilities of internet UGC and PGC resource libraries, as well as the large text model. This will be elaborated below:

[0129] Step 1. For the question "Which episode features Liu Huaqiang buying watermelons?", based on the intent recognition capability of the text big data model, the user intent is identified as a resource search intent. Then, based on the question, the search engine interface is called to request the search engine to search for relevant search results on the Internet.

[0130] Internet search results are presented in web page format and may include text results and / or image results.

[0131] Figure 4 shows several examples of internet search results (text results):

[0132] Search result 1: "[Source: Douyin Encyclopedia] Liu Huaqiang, male, is the Hengzhou gang leader in the 2003 Chinese mainland police and gangster TV series 'Conquest'..."

[0133] Search result 2. "[Source: Kuaikan Comics] In which episode of 'Conquest' does Liu Huaqiang buy watermelons? In episode 8 of the drama 'Conquest,' Sun Honglei plays a gangster boss named Liu Huaqiang who goes to a street stall to buy watermelons..."

[0134] Search result 3. "[Source: Douban.com 2024-04-19 13:49:15] In the TV series 'Conquest,' the scene where Liu Huaqiang buys watermelons appears in episode 9. Liu Huaqiang discovers something wrong with the scale while buying the watermelons..."

[0135] As can be seen from the examples of search results in Figure 4, although the internet can cover a wider and newer range of content, the content on the internet is not necessarily accurate or professional. For example, search result 2 indicates that the plot of "Liu Huaqiang buying watermelons" is in episode 8, while search result 3 indicates that the plot is in episode 9. Therefore, this application embodiment will combine professional, high-quality PGC (Professionally Generated Content) from the PGC resource library to perform resource searches, in order to provide users with accurate resource search results.

[0136] Step 2. Combine the question and internet search results into a first customized prompt. Use the first customized prompt to guide the text model to extract one or more video resource entities from the internet search results.

[0137] The essence of a video resource entity is information, which mainly refers to resource identification information such as the video resource name, and may also include other information such as the resource creator, resource type, and resource release time.

[0138] In the example in Figure 4, the video resource entities extracted from the Internet search results include "Episode 8 of 'Conquest'" and "Episode 9 of 'Conquest'", among others.

[0139] Step 2 utilizes the text-based large-scale model's ability to understand the intent and content of internet search results. The first customized prompt guides the text-based large-scale model to extract resource entities from internet search results.

[0140] The first customized prompt, as shown in Figure 5, may include, but is not limited to, the following parts: task description, entity notes, entity examples, search engine results, and question. The task description guides the text model on what task to perform; in the example in Figure 4, it specifically guides the text model to output video resource entities according to the entity notes and entity examples. The entity notes constrain the format in which the text model outputs the video resource entities. The entity examples provide examples to achieve better guidance. The search engine results require concatenating internet search results. The question requires concatenating the question "Which episode features Liu Huaqiang buying watermelons?".

[0141] Figure 5 is just one example. The first custom prompt can include different or more guiding items, and its format is not limited.

[0142] Step 3. Based on the video resource entity obtained in Step 2, search in the PGC resource library to obtain the PGC search results for the video resource entity.

[0143] The PGC search results for this video resource entity mainly include episode summaries, time-based summaries, and other film and television clips, and may also include film and television introductions, cast information, resource type, etc.

[0144] In the example in Figure 4, the PGC search results for the film and television entities "Episode 8 of 'Conquest'" and "Episode 9 of 'Conquest'" in the PGC resource library include: the time-sharing plot of Episode 8 of "Conquest" and the time-sharing plot of Episode 9 of "Conquest". Among them:

[0145] The plot of episode 8 of "Conquest" can be summarized as follows: "01:00-10:00 After learning that the police have preliminarily determined that Liu Huaqiang is a key suspect in the '4.28' case, Li Li immediately called Liu Huaqiang and told him to leave Hengzhou immediately..." ... "40:00-45:00 The cunning Liu Huaqiang has realized the danger approaching. He sent his mistress Li Mei to the hospital to investigate and found that Liu Huawen's ward had been searched. However, Liu Huaqiang did not stop and set his sights on the next target."

[0146] The time-sharing plot of episode 9 of "Conquest" is as follows: "01:00-10:00 After investigation, the police discovered the residence of Liu Huaqiang and others in Huadian Community, but the cunning Liu Huaqiang had already slipped away before the police found him...", ..., "30:00-40:00 When Liu Huaqiang was buying a watermelon, he asked the vendor, 'Is this watermelon guaranteed to be ripe?' Afterwards, Liu Huaqiang accused the vendor of tampering with the scale...", "40:00-46:37 The police launched a large-scale manhunt for Liu Huaqiang's subordinates, and a number of former members of Liu Huaqiang's gang were successively brought to justice."

[0147] Steps 4 to 7 mainly introduce the semantic understanding capabilities of text-based large-scale models and the image understanding capabilities of multimodal large-scale models, detailing the processes of filtering, summarizing, and providing resources for the search results obtained in steps 1 and 3. This will be elaborated below:

[0148] Step 4. Utilize the content understanding capabilities of the large text model to filter internet search results (UGC) and search results (PGC) in the PGC resource library based on the question, filtering out low-quality and inaccurate content, and determining the first video resource and the first video resource segment that match the user's intent.

[0149] For example, in the example in Figure 4, the internet search results for "Episode 8 of Conquest" indicated that the scene of "Liu Huaqiang buying watermelons" was in episode 8, but this was filtered out because it was inaccurate. The first video resource was episode 9 of Conquest, and the first video resource segment was the 30:00-40:00 segment of episode 9 of Conquest.

[0150] After identifying the first video resource, the terminal device can reply to the user with the video resource in the dialogue interface of the voice smart assistant, such as "In the TV series 'Conquest,' the classic scene of Liu Huaqiang buying melons appears in episode nine."

[0151] Furthermore, in step 4, based on the content of the filtered internet search results and PGC search results, the dialogue capabilities of the text big data model can be used to generate response text for the question, so as to reply to the user with the first video resource and the first video resource segment that match the user's intent. The response text may specifically include a segment introduction and a reason for recommendation, as shown in the example response text in Figure 4: "In the TV series 'Conquest,' the classic scene of Liu Huaqiang buying watermelons appears in episode nine. This episode is an important turning point in the drama. Liu Huaqiang (played by Sun Honglei) clashes with the watermelon vendor after discovering a magnet under the vendor's scale, and ultimately... This plot is widely remembered by the audience for its tense and exciting dramatic conflict and Sun Honglei's superb performance."

[0152] Step 5. Utilize the multimodal information processing capabilities of the multimodal big data model to analyze the filtered search results and the first video resource segment, and further locate the second video resource segment that better matches the user's intent from the first video resource segment, so as to obtain a more granular video segment that matches the user's intent.

[0153] In the example in Figure 4, the 30:00-40:00 segment from episode 9 of "Conquest" and related images from internet search results for episode 9 of "Conquest" can be input into a multimodal large model. After inference and calculation, the video resource segment that matches the user's intent is located around 32:51, so subsequent playback can start from 32:51. The 32:51-40:00 segment from episode 9 of "Conquest" is a more granular second video resource segment.

[0154] In the specific execution of step 5, to improve the accuracy of locating the second video resource segment, image results can be extracted from the filtered internet search results. These image results, along with the first video resource segment, are then input into the multimodal large model to obtain the second video resource segment. This is because text has a certain degree of ambiguity, while images are more explicit and specific. Images can help the multimodal large model more accurately locate the second video resource segment from the first video resource segment, such as using the position of the image within the first video resource segment as the starting frame position of the second video resource segment.

[0155] After the second video resource segment is determined, the server can transmit the indication information and the first frame image of the second video resource segment to the terminal device. In this way, the terminal device can use the indication information and the first frame image to render a resource service card, wherein the first frame image is used as the cover image of the resource service card.

[0156] Step 5 is a further improvement of the embodiments of this application. Its purpose is to more accurately locate video resource segments that better match the user's intentions and meet the user's fragmented viewing needs.

[0157] Step 6. Receive the response text from the server regarding the question. The terminal device can display the response text in the dialogue interface of the voice assistant.

[0158] Step 7. Receive the instruction information and first frame image of the second video resource segment transmitted by the server. The terminal device can display the video resource service card in the dialogue interface of the voice smart assistant, such as the service card for the segment 32:51-40:00 in episode 9 of "Conquest". The first frame image can be used as the cover image of the resource service card.

[0159] Upon detecting that a user has clicked on the resource service card, the terminal device can open the video playback interface and automatically play the second video resource clip. Furthermore, the user can also trigger the playback of the second video resource clip by continuing the conversation (e.g., saying "Thank you, please play").

[0160] The technical effects of Example 1 include at least the following: by leveraging the abundant UGC and high-quality PGC in the video resource library on the internet, as well as the semantic and content understanding capabilities of the text big data model, users can accurately search for the video resources they want through conversational questions. Furthermore, after identifying video resource segments that match the user's intent using internet search results and PGC search results, the multimodal big data model can further locate more granular segments within those video resource segments that match the user's intent, thus meeting the user's fragmented viewing needs.

[0161] Example 2

[0162] Example 2 is mainly used to handle users' fuzzy search or recommendation problems. It combines user profiles, Internet search results and PGC search results to automatically mine resources that match user intent using a large text model.

[0163] Figure 6 illustrates the overall flow of the resource search method provided in Embodiment 2. The details are as follows.

[0164] S21. The terminal device can receive a question uttered by the user and transmit the question to the server.

[0165] S22. Based on this question, the server identifies the user's intent as a resource search intent, and searches the Internet based on this question, obtaining multiple Internet search results.

[0166] S23. The server can extract one or more resource entities from multiple Internet search results.

[0167] S24. The server can use the resource entity obtained in S23 to search in the PGC resource library and obtain the PGC search results for that resource entity.

[0168] For detailed explanations of S21-S24, please refer to S11-S14 in Implementation 1, which will not be repeated here.

[0169] After receiving the question, the terminal device can also send a request to the user profiling platform to obtain the user profile and transmit the user profile to the server.

[0170] S25. The server can analyze internet search results and PGC search results based on user profiles and large text models to obtain the optimal resource features that match the user profile.

[0171] The optimal resource feature describes which resources match the user's resource preferences. It can include, but is not limited to, one or more of the following parameters: optimal resource entity, optimal fragment feature, optimal person feature, etc. In this paper, "optimal" refers to the result output by the large text model, not the absolute optimal one. There may be resource features that match the user's resource preferences better than the optimal one, and the degree of optimization depends on the capabilities of the large model.

[0172] In the specific implementation, the server can execute S25 using a second customized prompt and a large text model. The second customized prompt can be used to guide the large text model to analyze the optimal resource features based on the question, user profile, the aforementioned internet search results, and the aforementioned PGC search results. The implementation of the second customized prompt will be explained in detail below, and will not be elaborated on here.

[0173] S26. Based on the optimal resource characteristics obtained in S25 and the aforementioned Internet search results and PGC search results, the server can further locate recommended resource fragments from the optimal resources.

[0174] The optimal resource refers to the resource represented by the optimal resource entity.

[0175] In specific implementation, the multimodal large model can be used to execute S26: input the aforementioned Internet search results and the aforementioned PGC search results, as well as the optimal resource features, into the multimodal large model, and obtain the resource fragments that match the optimal resource features, i.e., the recommended resource fragments, through the inference operation of the multimodal large model.

[0176] Thus, different users often receive different recommended resource snippets, even if they ask the same question. Assuming the first user and the second user are different, even if they both ask the voice assistant the same question, such as "What are some good variety shows lately?", the recommended resource snippets output for the first user and the second user will be different, and the reply text and resource service cards they receive in the dialogue interface will also be different. User identity can be determined based on the user's logged-in account, or it can be determined based on other user characteristic data such as the user's voice characteristics (the voice characteristics collected when the question is spoken).

[0177] S27. The server can utilize the semantic and content understanding capabilities of the large text model to comprehensively analyze the aforementioned internet search results, PGC search results, and optimal resource features to generate a response text addressing the question. The server can then transmit this response text to the terminal device, allowing the terminal device to output the response text to the user within the voice assistant's dialogue interface. The terminal device can also read the response text aloud.

[0178] The response text can be used to introduce the recommended resource snippets, which may include a brief introduction, reasons for recommendation, etc., so that users can better understand the best resources, recommended resource snippets, and reasons for recommendation.

[0179] For example, in response to the question "Which episode features the blood test in 'The Legend of Zhen Huan'?", the reply could be: "'The Legend of Zhen Huan' is a TV series, not a movie. The blood test scene occurs in episode 63. In this episode, Zhen Huan is accused by Consort Qi of having an improper relationship with Imperial Physician Wen, which leads to..."

[0180] When generating recommendation reasons and summary introductions, internet search results help to broaden the scope of content, while PGC search results help to improve the accuracy of content.

[0181] Furthermore, the server can also transmit indication information for recommended resource clips to the terminal device, such as the title of the film or television work, the number of episodes, and the time range of the video clip. The terminal device can then render a resource service card based on this indication information and display it in the dialogue interface of the voice assistant. Moreover, when a user triggers playback of a recommended resource clip, the terminal device can use the indication information to send a request to the video application's server to retrieve the recommended resource clip and play it on the device.

[0182] In addition, the server can transmit introductory images of recommended resource snippets to the terminal device. The terminal device can then use these images as the cover image for the resource service card of the recommended resource snippet, thereby increasing the attractiveness of the resource service card to the user. Recommended resource snippets can be extracted by the server from the best resources.

[0183] The resource search method provided in Embodiment 2 will be further explained in detail below using video resources as an example, through the example in Figure 7.

[0184] In the example of Figure 7, the question is "What are some good variety shows lately?" A personalized video clip supply solution for this question could include the following steps:

[0185] Steps 1 to 3 mainly introduce the process of searching for resources based on the intent recognition and content understanding capabilities of the internet UGC and PGC resource libraries, as well as the large text model. For details of these steps, please refer to the specific explanation of steps 1 to 3 in Figure 4, which will not be repeated here.

[0186] The additional step compared to the example in Figure 4 is that, when identifying user intent based on questions, user profiles can also be obtained from the user profiling platform. The example in Figure 7 illustrates user profiles based on aspects such as gender, age, occupation, movie preferences, and actor preferences.

[0187] In the example in Figure 7, the video resource entities extracted from internet search results include "Chinese Restaurant Season 8", "Happy Night", and "Escape Room Season 6". The PGC search results for each of these video resource entities from the PGC resource library can include: episode summaries, time-sharing episode summaries, and other film and television clips, as well as: variety show type, guests, video platform, and other content.

[0188] Step 4. Concatenate the question, user profile, the aforementioned internet search results, and the aforementioned PGC search results into the second customized prompt. Use the second customized prompt to guide the text big model to output the optimal resource features based on the user profile, the aforementioned internet search results, and the aforementioned PGC search results.

[0189] Optimal resource features enable subsequent multimodal large models to locate video resource segments from optimal video resources that also match user preferences.

[0190] In the example of Figure 7, the optimal resource features include: optimal video resource, optimal segment feature, and optimal person feature. Among them:

[0191] The optimal video resource is "Escape Room Season 6" because: the guest is Da Zhangwei, which matches the user's preferred actors "Da Zhangwei, Yuan Yawei..."; and the variety show type is "educational real-life puzzle interactive variety show", which also matches the user's viewing preferences "exciting, quirky, puzzle-solving...".

[0192] The optimal segment characteristics are "exciting scenes, adventure sequences, puzzle-solving scenarios..." because these match the user's viewing preferences.

[0193] The optimal character traits are "Da Zhangwei, singer, comedian" because this matches the user's preference for actors.

[0194] The second customized prompt, as shown in Figure 8, may include, but is not limited to, the following parts: task description, answer rules, knowledge base, and question. The task description guides the text-based big data model on what task to perform. In the example in Figure 7, it specifically guides the model to integrate the knowledge base and the user's personalized profile to output video resource entities that match the user's intent and the optimal video resource features. The answer rules guide the model on how to recommend video resources based on the knowledge base and can prevent fabrication to avoid the "big model illusion" problem. The knowledge base needs to combine the aforementioned internet search results and PGC search results. The "search engine results are as follows" section can combine internet search results, and the "video resource library" section can combine PGC search results. The question needs to be "What are some good variety shows recently?". Furthermore, the user profile can also be incorporated into the second customized prompt.

[0195] Figure 8 is just one example. The second custom prompt can include different or more guiding items, and its format is not limited.

[0196] Step 5. Input the optimal resource features obtained in Step 4, the aforementioned Internet search results, and the aforementioned PGC search results into the multimodal large model to locate the recommended film and television segments in the optimal video resources.

[0197] Recommended movie clips are more granular video resources that are matched with user profiles, satisfying users' fragmented viewing needs while also meeting their personalized viewing preferences.

[0198] In the example in Figure 7, recommended film clips include: "Episode 1 of Escape Room Season 6, 03:45-10:00", "Episode 2 of Escape Room Season 6, 25:12-31:28", and "Episode 3 of Escape Room Season 6, 42:36-49:01".

[0199] After locating the recommended movie clip from the optimal video resources, the server can transmit the instruction information and the first frame image of the recommended movie clip to the terminal device. In this way, the terminal device can use the instruction information and the first frame image to render a resource service card, where the first frame image is used as the cover image of the resource service card.

[0200] Step 6. Receive the response text for the question transmitted by the server. The terminal device can display the response text for the question in the dialogue interface of the voice assistant.

[0201] The response text can be generated based on the optimal resource features obtained in step 4 and the recommended film clips obtained in step 5, in order to recommend video resource clips that match the user's intent and preferences to the user. Specifically, the response text may include introductions to each recommended resource clip, reasons for recommendation, etc., such as: "In the first episode of 'Escape Room Season 6' from 03:45-10:00, Da Zhangwei's analysis of the tram puzzle is brilliant and insightful, earning praise from the other members..."; "In the second episode of 'Escape Room Season 6' from 25:12-31:28, Da Zhangwei faces his fear and challenges his personal tasks; will he be able to pass the test..."; "In the third episode of 'Escape Room Season 6' from 42:36-49:01, facing the pursuit of zombies, Da Zhangwei bravely confronts them, protecting the team, and the team spirit erupts...".

[0202] In addition, upon receiving the instruction information and first frame image of recommended movie clips transmitted from the server, the terminal device can also display video resource service cards in the dialogue interface of the voice assistant. The first frame image can be used as the cover image of the resource service card. When the user clicks on the resource service card, the terminal device can open the movie playback interface and automatically play the recommended movie clip. If there are multiple recommended movie clips, the terminal device can display the service cards for each clip in the dialogue interface. The user can select and click on a specific resource service card to jump to the playback interface of the corresponding recommended movie clip, where that clip will play automatically.

[0203] The technical effects of Example 2 include at least the following: based on user profiles, large text models, and abundant internet content and high-quality PGC content, it supports users in finding optimal resources that match their intentions and preferences through fuzzy search or recommendation questions. Furthermore, based on the multimodal large model's ability to understand various types of content such as text, images, and videos, it can accurately locate the recommended resource segments that users are most interested in from the optimal resources, meeting users' needs for fragmented viewing and achieving the goal of accurate resource recommendation.

[0204] In Example 2, the user profile further includes real-time user characteristics, such as the user's current environment, surrounding people, specific holidays, and weather. This allows the large text model to further combine these real-time user characteristics to recommend resources or resource snippets, resulting in a more immersive and user-centric resource search experience. These real-time user characteristics can be obtained through the terminal device's perception of the physical world.

[0205] Example 3

[0206] Example 3 focuses on how to support a user in selecting resources by continuing the conversation, provided that the user retrieved multiple resources in the previous round of dialogue. Example 3's solution occurs in a multi-turn dialogue scenario. The user's fuzzy search question or recommendation question can trigger this multi-turn dialogue.

[0207] Figure 9 illustrates the overall flow of the resource search method provided in Embodiment 3. The details are as follows.

[0208] First round of dialogue (S31-S36)

[0209] S31. In the first round of dialogue, the terminal device can receive a question, such as "Recommend three movies suitable for family reunion", and transmit the question to the server.

[0210] S32. Based on this question, the server identifies the user's intent as a resource search intent, and searches the Internet based on this question, obtaining multiple Internet search results.

[0211] S33. The server can extract resource entities from multiple internet search results.

[0212] S34. The server can use the resource entity obtained in S33 to search in the PGC resource library and obtain the PGC search results for that resource entity.

[0213] For a detailed explanation of S31-S34, please refer to S11-S14 in Implementation 1, which will not be repeated here.

[0214] S35. The server can utilize the content understanding capabilities of the large text model to first filter the aforementioned internet search results and the aforementioned PGC search results, filtering out low-quality and inaccurate content. Then, based on the question, the filtered internet search results, and the PGC search results, it can determine multiple resources that match the user's intent and / or multiple resource fragments from these multiple resources, and generate a response text.

[0215] Correspondingly, the server can also transmit indication information of these resources and / or resource fragments to the terminal device, such as the name of the film or television work, the number of episodes, and the time range of the video clip. In this way, the terminal device can render the corresponding resource service cards based on the indication information of these resources and / or resource fragments, and display these resource service cards in the dialogue interface of the voice intelligent assistant.

[0216] Furthermore, the server can also transmit the reply text to the terminal device, so that the terminal device can display the reply text in the dialogue interface of the voice assistant.

[0217] In addition, the server can also transmit introductory images of these resources and / or resource fragments to the terminal device. The terminal device can then use these introductory images as cover images for the corresponding resource service cards, thereby enhancing the appeal of the resource service cards to users. Specifically, the introductory image for a resource can come from the resource itself, such as the first frame image; the introductory image for a resource fragment can come from the resource fragment, such as the first frame image.

[0218] S36. Based on the indication information of the multiple resources and / or resource segments determined in S35, the terminal device can also display resource service cards for each of these multiple resources and / or resource segments that match the user's intent in the dialogue interface of the voice intelligent assistant. The user can click on a resource service card to view the automatic playback of the corresponding resource or resource segment.

[0219] In Example 3, the problem is a fuzzy search problem or a recommendation problem. Therefore, the large text model often identifies multiple resources that match the user's intent. Thus, the response text can include an introduction and reasons for recommendation for each resource, facilitating user reference for subsequent resource selection. There are also multiple resource service cards; users can click on a specific card to jump to the playback interface of the corresponding resource or resource segment and watch its automatic playback.

[0220] In addition to manually selecting resource service cards, users can also select resources by continuing the conversation. The technical implementation will be described in detail below.

[0221] Second round of dialogue (S37-S40)

[0222] S37. After updating the dialogue interface via S35 and S36, the terminal device can ask the user which resource to select through the voice assistant.

[0223] For example, a question text, such as "Find these videos, which one do you want to watch?", can be appended to the end of the aforementioned reply text. The voice assistant can also read this question text aloud.

[0224] S38. In the second round of dialogue, the terminal device may receive a question, such as "Play the first one", and transmit the question to the server.

[0225] S39. The server can identify the user's intent as a resource selection intent based on this question, and use a large text model to analyze information such as the aforementioned internet search results, the aforementioned PGC search results, the previous round of dialogue text, the rendered content of the resource service cards, and the position of the resource service cards, to select the target resource and / or target resource fragment that matches the user's intent from multiple resources and / or resource fragments. Here, the position of the resource service cards refers to their arrangement order.

[0226] Correspondingly, the server can also transmit indication information of the target resource and / or target resource segment to the terminal device, such as the name of the film or television work, the number of episodes, and the time range of the video segment. In this way, the terminal device can request the target resource and / or target resource segment from the server of the video application based on the indication information of the target resource and / or target resource segment, and play the target resource and / or target resource segment on the terminal side.

[0227] The previous round of dialogue text may include the aforementioned response text, which describes the introduction and reasons for recommendation of each resource and / or each resource fragment. Users may use certain descriptions that appeared in the previous round of dialogue text to make resource selections. Users may also use other content such as plot, actors, and creation time to describe resource selections. Users may also utilize the location of resource service cards and rendered content in the dialogue interface to describe resource selections. This application embodiment fully utilizes the intent recognition and semantic understanding capabilities of the large text model to support users in describing resource selections in these ways, enabling users to express their resource selections more freely.

[0228] In practice, the terminal device can execute S40 using a third-party customized prompt and a large text model. The third-party customized prompt can be used to guide the large text model to select a target resource or target resource fragment that matches the user's intent from multiple resources and / or multiple resource fragments pushed to the user in the previous round of dialogue. The implementation of the third-party customized prompt will be explained in detail below, and will not be elaborated on here.

[0229] S40. Based on the indication information of the target resource and / or target resource segment transmitted by the server, the terminal device may request the server of the video application to obtain the target resource and / or target resource segment, and play the target resource and / or target resource segment on the terminal side.

[0230] The resource search method provided in Embodiment 3 will be further explained in detail below using video resources as an example, with reference to Figure 10.

[0231] In the example in Figure 10, the question in the first round of dialogue is "Play the movie where Andy Lau plays a policeman," and the question in the second round is "Play the one with Tony Leung." Solutions for supporting user selection of video resources in multi-turn dialogue scenarios may include the following steps:

[0232] Steps 1 to 3 mainly introduce the process of searching for resources based on the intent recognition and content understanding capabilities of the internet UGC and PGC resource libraries, as well as the large text model. For details of these steps, please refer to the specific explanation of steps 1 to 3 in Figure 4, which will not be repeated here.

[0233] Step 4. The server can utilize the content understanding capabilities of the large text model to filter internet search results (UGC) and search results (PGC) in the PGC resource library for the question "Play the movie that Andy Lau plays as a policeman". Low-quality and inaccurate content will be filtered out. Based on the question, the filtered internet search results and PGC search results, multiple video resources that match the user's intent and / or multiple video resource segments from these multiple video resources will be determined, and a response text will be generated.

[0234] Correspondingly, the server can also transmit indication information of these video resources and / or video resource segments to the terminal device, such as the name of the film or television work, the number of episodes, and the time range of the video segment. In this way, the terminal device can render the corresponding resource service cards based on the indication information of these video resources and / or video resource segments, and display these resource service cards in the dialogue interface of the voice intelligent assistant.

[0235] Furthermore, the server can also transmit the reply text to the terminal device, so that the terminal device can display the reply text in the dialogue interface of the voice assistant.

[0236] In the example in Figure 10, the multiple video resources that match the user's intent are: "Shock Wave 2", "Infernal Affairs", "Chasing the Dragon", "Shock Wave", "Cold War", "Storm", and "Future Police".

[0237] In the example in Figure 10, the reply text includes introductions and reasons for recommendation for each of these video resources, such as: "1. *Shock Wave 2*: Andy Lau plays a bomb disposal expert in the film...; 2. *Infernal Affairs*: Andy Lau plays undercover police officer Lau Kin-ming in this film...; 3. *Chasing the Dragon*: Andy Lau plays Detective Lui Lok in this film, alongside Donnie Yen...; 4. *Shock Wave*: Andy Lau plays the lead role of Cheung Choi-shan in this film, a member of the Hong Kong Police Force...; 5. *Cold War*: Although Andy Lau is not a main character in this film, as part of the Hong Kong Police Force, the Asian financial center...; In addition, Andy Lau has also played police roles in other films, such as Senior Inspector Lui Ming-tsan in *Firestorm*, and a police officer from the future in *Future X-Cops*,..."

[0238] Step 5. Based on the indication information of the multiple video resources and / or video resource segments determined in Step 4, the terminal device can also display the resource service cards of the multiple video resources and / or video resource segments in the dialogue interface of the voice smart assistant.

[0239] After obtaining the multiple video resources and / or video resource clips, users can select resources by continuing the conversation, such as saying "Play the one with Tony Leung".

[0240] Step 6. The server can concatenate the question from the current dialogue ("Play the one with Tony Leung"), the text from the previous dialogue, and the rendered content of the resource service card into a third-party customized prompt. Using this third-party customized prompt, the server guides the text model to select the target video resource or video resource segment that matches the user's intent from multiple video resources and / or video resource segments pushed to the user in the previous dialogue. In this way, the user can use the content presented in the dialogue interface to describe their intended video resource or video resource segment.

[0241] In addition, the terminal device can also splice the aforementioned Internet search results, PGC search results, and the location of resource service cards into a third-party customized prompt, so as to support users to use content outside the dialog interface (such as a certain movie plot) and the location of video resources in the screen to describe the resources they intend to select, providing more freedom of expression.

[0242] Step 7. After selecting the target video resource or target video resource segment, the server can also return indication information of the target video resource or target video resource segment to the terminal device to trigger the terminal device to play the target video resource or target video resource segment. Accordingly, after receiving the indication information, the terminal device can use the indication information to send a request to the server of the video application to request the acquisition of the target video resource or target video resource segment, and play the target video resource or target video resource segment on the terminal side.

[0243] In the example in Figure 10, the selected target video resource is "Infernal Affairs". The text big data model comprehensively analyzes the response text of the first round of dialogue, "Infernal Affairs: In this film, Andy Lau plays an undercover police officer, Lau Kin-wu, who infiltrates a triad and swaps identities with a police academy student played by Tony Leung, unfolding a dual narrative," the search results in the PGC resource library, "Infernal Affairs is a Hong Kong film released in 2002, produced by Media Asia Films, directed by Andrew Lau and Alan Mak, starring Andy Lau and Tony Leung...", and the internet search results, "Infernal Affairs is a Hong Kong film released in 2003, produced by Media Asia Films, directed by Andrew Lau and Alan Mak, starring Andy Lau and Tony Leung...", and combines this with the question "Play the one with Tony Leung", to determine that the video resource the user intends to select is "Infernal Affairs".

[0244] The third customized prompt, as shown in Figure 11, may include, but is not limited to, the following parts: task description, answer rules, constraints, knowledge base, and questions. The task description guides the text model on what task to perform. In the example in Figure 11, it specifically guides the text model to decide whether to play or which video resource to play based on the new question, given the previous dialogue and the video resources pushed to the user. The answer rules guide how the text model responds to different user intentions. The constraints guide the text model to correct repeated user expressions, such as repeated sequence numbers. The knowledge base requires concatenating the previous dialogue text and card rendering content. The "Text reply content is as follows" section can concatenate the previous dialogue text (including introductions and recommendation reasons for each video resource), and the "Card rendering content is as follows" section can concatenate the card rendering content for each video resource. The question section concatenates the question from the current dialogue (e.g., "Play the one with Tony Leung").

[0245] Figure 11 is just one example. The third custom prompt can include different or more guiding items, and its format is not limited.

[0246] Figure 10 does not illustrate the supply of resource segments because the user's intent expressed in the multi-turn dialogue is to select a movie, not a segment from a movie. In practical applications, the user's intent expressed in multi-turn dialogues can also be the intent to select a video resource segment. For example, in the first round, a user says "Play the movie where Andy Lau plays a policeman," and after receiving several related movies, the user says in the second round, "Play the rooftop showdown scene from 'Infernal Affairs,'" to express the intent to select a resource segment. As another example, in the first round, a user says "Recommend several videos of police catching thieves," and after receiving several related recommended videos, the user says in the second round, "Play the first one," to express the intent to select a resource segment.

[0247] In the example of Figure 10, even if the user does not explicitly express the intention to select a video resource segment to play, this embodiment of the application can also locate the video resource segment that matches the user's preferences from the target video based on the user profile, and push the instruction information of the video resource segment to the terminal device. In this way, the terminal device can push the video resource segment to the user, such as rendering it as a video resource card and pushing it to the customer.

[0248] The technical effects of Embodiment 3 include at least the following: in scenarios involving fuzzy search or resource recommendation, multiple resources can be provided to the user, along with resource descriptions and reasons for recommendation, making it easy for the user to quickly understand each resource and its reasons for recommendation; moreover, users can select the resource they want to view from multiple resources using any natural language, without the need for manual selection; from resource query and selection to resource playback control, the entire process can be completed through spoken dialogue, which is simple and efficient.

[0249] Based on the above embodiments, the resource search method provided in this application is summarized below. Its overall flow is shown in Figure 12, including:

[0250] S51. The terminal device receives the first dialogue content input by the first user, and the first dialogue content expresses the intention to search for resources in natural language.

[0251] Users can input the initial dialogue content through a smart voice assistant, such as by speaking the initial dialogue content aloud. However, users can also input the initial dialogue content via text input; this application embodiment does not limit the input method for the initial dialogue content.

[0252] S52. Based on the content of the first dialogue, the terminal device may reply to the first user with a first resource fragment that matches the resource search intent, wherein the first resource fragment is a video resource fragment.

[0253] The first resource segment can be one or more episodes of a film or television work, or one or more segments from one episode of a film or television work. For example, the former could be episode 63 of "Empresses in the Palace," and the latter could be the segment of "Zhen Huan's Blood Test of Kinship" in episode 63 of "Empresses in the Palace."

[0254] The process of replying to the first user with a first resource fragment that matches their search intent can include one or more of the following methods: displaying a first reply text or displaying a first service card. The first reply text can describe the first resource fragment and may include reasons for recommending the fragment, a plot summary, etc. The first reply text can be, for example, reply text 15 in Figure 1A or reply text 18 in Figure 1B. The first service card can be used by the user to view the first resource fragment. The first service card can be, for example, service card 16 in Figure 1A or service card 19 in Figure 1B.

[0255] The rendered content on the first service card can also serve to describe the first resource fragment. Its rendered content may include one or more of the following: the name, type, release time, creator, cover image, etc. of the first resource fragment or its associated video resource.

[0256] As shown in S53, when the terminal device detects a user operation applied to the first service card, the terminal device can jump to the playback interface of the first resource segment and play the first resource segment. Besides clicking the first service card as user input, the terminal device can also receive user input to trigger the playback of the first resource segment through a continuation dialogue, and then jump to the playback interface of the first resource segment to play it. The continuation dialogue may include: first asking whether to play the first resource segment; then receiving the user's response to the inquiry. This continuation dialogue can be a voice dialogue or a text dialogue; this embodiment does not limit the dialogue format. The implementation details of triggering the playback of the first resource segment can be found in the relevant content of the foregoing embodiments, and will not be repeated here.

[0257] Furthermore, different users can receive different recommended resource fragments, even if different users input the same dialogue content, which may be a fuzzy search question or a recommendation question. Specifically, the terminal device can receive the first dialogue content input by the second user; based on the first dialogue content input by the second user, the terminal device can reply to the second user with a second resource fragment that matches the resource search intent. The second user is different from the first user, and the second resource fragment is different from the first resource fragment. The second resource fragment can be one or more episodes of a film or television work, or one or more clips from an episode of a film or television work. Replying to the second user with a second resource fragment that matches the resource search intent can include one or more of the following methods: displaying a second reply text, displaying a second service card. The second reply text can be used to describe the second resource fragment, including reasons for recommending the second resource fragment, plot summaries, etc. The second service card can be used by the user to view the second resource fragment.

[0258] The memory information or personal profiles of the first user and the second user on the terminal device may differ. Memory information may include the user's video browsing history, video collection history, and other usage traces that can indicate the user's preferences for film and television resources. The personal profile may be a user profile determined by the cloud side based on user data uploaded from the terminal (such as gender, age, occupation, etc.) and / or the aforementioned memory information, used to indicate the user's film and television preferences. User film and television preferences can be broad, including preferences for film and television content, preferences for film and television creators (such as actors), preferences for film and television sources, preferences for playback resolution, etc.

[0259] Based on different users' film and television resource preferences, this application embodiment recommends film and television resources that match user preferences in a targeted manner. The second resource segment differs from the first resource segment in the following ways, but is not limited to: the video content of the second resource segment is different from that of the first resource segment; or the content of the second resource segment is the same as that of the first resource segment but the resource to which they belong is different; or the content of the second resource segment is the same as that of the first resource segment but the clarity is different.

[0260] Furthermore, there may be multiple first resource segments that meet the resource search intention. In this case, after replying to the first user with the first resource segment that meets the resource search intention, the terminal device can further receive the second conversation content input by the first user. The second conversation content expresses the resource selection intention of selecting the third resource segment from multiple first resource segments in natural language. Then, the terminal device can select the third resource segment from multiple first resource segments according to the second conversation content and jump to the playback interface of the third resource segment to play the third resource segment.

[0261] The first resource segment that meets the resource search intention can be determined by the terminal device according to the first conversation content, especially when the terminal device has strong computing power. The first resource segment that meets the resource search intention can also be determined by the cloud server according to the first conversation content, and the first resource segment is fed back to the terminal device, and the terminal device replies to the user with the first resource segment. The latter method is the end-cloud collaboration method.

[0262] The following describes an implementation method of end-cloud collaboration with reference to FIG. 13:

[0263] As shown in S61, S64 and S65, before the terminal device replies to the user with the first resource segment that meets the resource search intention: the terminal device can send the first conversation content to the server. Correspondingly, the server can receive the first conversation content sent by the terminal device and determine the first resource segment that meets the resource search intention for the first conversation content. Then, the server can return the indication information of the first resource segment to the terminal device. In this way, the terminal device can receive the indication information of the first resource segment from the server, and then know which resource segment meets the user's search intention, and execute S52 to feedback the search result to the user.

[0264] Among them, the indication information of the first resource segment may include: the identification information of the resource to which the first resource segment belongs, and the positioning information of the first resource segment in its所属 resource. The identification information of the resource may be, for example, the resource name, and the positioning information of the first resource segment in its所属 resource may be the resource segment number, time range, etc. For example, the indication information of the first resource segment is composed of the resource name of "Empresses in the Palace" and the resource segment number of "Episode 63". Another example is that the indication information of the first resource segment is composed of the resource name of "Infernal Affairs" and the time range of "30:00-40:00". When using the time range as the positioning information of the first resource segment in its所属 resource, the start time of the first resource segment can be used as the key parameter indicating the time range, and the end time of the first resource segment can be defaulted to the end time of the resource to which the first resource segment belongs if not clearly indicated.

[0265] Specifically, the server can utilize the interconnected search results and PGC search results associated with the first dialogue content to determine the first resource fragment that matches the resource search intent. As shown in S62a and S62b, the internet search results can be obtained by searching the internet based on the first dialogue content; as shown in S63a and S63b, the PGC search results can be obtained by searching the PGC resource library based on the first resource identifier. The PGC resource library records the resource fragments included in different resources. The first resource identifier belongs to the video resource entity mentioned above, and it can be extracted from the interconnected search results. Specifically, the server can use a first customized prompt guiding text model to extract resource identifiers such as video resource names from the internet search results. Regarding the implementation of the first customized prompt, please refer to the relevant content in the preceding embodiments, which will not be repeated here.

[0266] Furthermore, compared to the resource fragments recorded in the PGC resource library, the first resource fragment can be a shorter fragment, meaning it can be a more granular fragment. To this end, the server can first determine a third resource fragment that matches the user's search intent based on internet search results and PGC search results. The duration of the third resource fragment can be the same as the resource fragments recorded in the PGC resource library. Then, the server can use a multimodal large model to process the third resource fragment, internet search results, and PGC search results to further locate the first resource fragment from the third resource fragment. In this way, the first resource fragment can meet the user's fragmented viewing needs.

[0267] As shown in S67, before the terminal device replies to the user with a first resource fragment that matches the resource search intent: in addition to the indication information of the first resource fragment, the server can also return a first reply text to the terminal device. Accordingly, the terminal device can receive the first reply text. In this way, the terminal device can display the first reply text to present the user with an introduction to the first resource fragment and the reasons for its recommendation, allowing the user to clearly understand the first resource fragment and the recommendation logic.

[0268] Specifically, as shown in S66, the server can use a large text model to analyze the internet search results and PGC search results related to the first dialogue content, and generate the first response text. This can be referred to in the relevant content of the preceding embodiments, and will not be repeated here.

[0269] As shown in S68, before the terminal device replies to the user with a first resource fragment that matches the resource search intent: in addition to the indication information of the first resource fragment, the server can also return a first image to the terminal device. The first image can be generated based on the first resource fragment. Specifically, the first image can come from the first resource fragment, such as the first frame image of the first resource fragment. In this way, the terminal device can use the first image as the cover image of the first service card when rendering and generating the first service card, thereby increasing the user appeal of the first service card.

[0270] As shown in S69, after S67-S68, the terminal device can use the first reply text and the first image to render and generate a first service card, and reply to the user with the first resource fragment by displaying the first service card and / or the first reply text.

[0271] In this embodiment of the application, in order to improve the accuracy of the response, before the server determines the first resource fragment that matches the resource search intent based on the first dialogue content, the server can first use PGC search results to filter out inaccurate content in Internet search results.

[0272] Furthermore, recommended resource snippets can also be based on user profiles. This way, different users can receive different recommended resource snippets, even if different users input the same question—whether it's a fuzzy search question or a recommendation question.

[0273] Specifically, when recommending the first resource fragment to the first user: the server can analyze internet search results and PGC search results related to the first dialogue content based on the first user's user profile to determine the first resource fragment. In specific implementation, the server can first use a second customized prompt guiding text model to output optimal resource features based on the user profile, internet search results, and PGC search results; then, based on the optimal resource features and the internet search results and PGC search results, the server can further locate the first resource fragment from the optimal resources to address the first user's question. The optimal resource features include the identification information of the optimal resource and the optimal fragment features.

[0274] For the implementation of the second customized prompt, please refer to the relevant content in the aforementioned embodiments, which will not be repeated here.

[0275] There may be multiple first resource fragments that match the resource search intent. In this case, after replying to the first user with the first resource fragments that match the resource search intent, the terminal device can further receive second dialogue content input by the first user. The second dialogue content expresses the resource selection intent of choosing a third resource fragment from the multiple first resource fragments in natural language. Then, the terminal device can select the third resource fragment from the multiple first resource fragments according to the second dialogue content and jump to the playback interface of the third resource fragment to play the third resource fragment.

[0276] There may be multiple first resource fragments that match the resource search intent. In this case, as shown in S70-S72, the server can also receive second dialogue content sent by the terminal device. The second dialogue content expresses, in natural language, the resource selection intent to select a third resource fragment from the multiple first resource fragments. Then, the server can determine the third resource fragment that matches the resource selection intent from the multiple first resource fragments based on the second dialogue content and return indication information of the third resource fragment to the terminal device. The indication information of the third resource fragment can be used to indicate the third resource fragment among the multiple first resource fragments.

[0277] Specifically, the server can use a third prompt to guide the text model to select a third resource fragment that matches the user's resource search intent from multiple first resource fragments pushed to the user in the previous round of dialogue. The implementation of the third customized prompt will be explained in detail below, and will not be elaborated upon here. For details on the implementation of the third customized prompt, please refer to the relevant content in the aforementioned embodiments; it will not be repeated here.

[0278] Figure 14 illustrates, exemplarily, a terminal device 300 provided in an embodiment of this application.

[0279] Terminal device 300 can possess both human-computer interaction capabilities and computing capabilities. The device type of terminal device 300 can be any of the following: mobile phone, tablet computer, in-vehicle device (also known as vehicle infotainment system), smart home device such as smart large screen, wearable device such as smartwatch and smart glasses, extended reality (XR) device such as augmented reality (AR), virtual reality (VR), and mixed reality (MR), smart city device, etc.

[0280] As shown in Figure 14, the terminal device 300 may include: a processor 110, a memory 120, a display 130, a display driver integrated circuit (DDIC) 140, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a gyroscope sensor 180B, an accelerometer sensor 180E, and a touch sensor 180K, etc. The various components of the terminal device 300 can be connected via a bus.

[0281] The processor 110 provides computing power and can be used as the computing module of the terminal device 300. The display 130, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, buttons 190, motor 191, indicator 192, camera 193, and other input / output components provide human-computer interaction capabilities and can be used as the human-computer interaction module of the terminal device 300. When the computing module within the terminal device 300 has powerful computing capabilities, the terminal device 300 can independently execute the resource search method provided in this application embodiment. When the computing module within the terminal device 300 does not have powerful computing capabilities, the terminal device 300 can also only execute the human-computer interaction steps in the resource search provided in this application embodiment, such as receiving a query through a voice assistant, outputting a reply text, displaying resource service cards, and playing / jumping, etc., while the inference calculation steps based on the text-based large model and multimodal large model in this method can be executed by a cloud-based server with stronger computing power.

[0282] Processors 110 can be one or more, and they can be integrated into an integrated circuit of a system-on-a-chip (SOC). An SOC is a system-on-a-chip. Processors 110 may include a central processing unit (CPU), a graphics processing unit (GPU), and a neural network processing unit (NPU). The CPU may include an application processor (AP) and a baseband processor (BP). The AP is responsible for running the operating system, user interface, and applications on the terminal device 300; the BP is responsible for transmitting and receiving wireless signals and managing radio frequency services. The GPU is responsible for graphics rendering, performing tasks such as shading, material filling, rendering, and output based on rendering instructions and data from the CPU. The NPU, by referencing biological neural network structures, such as the transmission patterns between neurons in the human brain, can quickly process input information and continuously learn. The NPU can be used to run artificial intelligence algorithms, such as instruction recommendation algorithms, image processing algorithms, and image understanding algorithms. The CPU and GPU can be used to render and synthesize the image to be displayed on the monitor 130.

[0283] The memory 120 may include a program storage area and a user data storage area. The program storage area may store the operating system and one or more applications (such as games), while the data storage area may store data created by the user during use of the terminal device 300 (such as photos and contacts). The memory 120 may be a high-speed random access memory or a non-volatile memory, such as a hard disk, flash memory, or universal flash storage (UFS). The memory 120 may also be an external memory card, such as a Micro SD card.

[0284] The memory 120 may also store code instructions for the resource search method provided in the embodiments of this application. When the processor 110 reads the code instructions from the memory 120 and runs the code instructions, the terminal device 300 may execute the steps performed by the human-computer interaction module and / or the computing module in the resource search method provided in the embodiments of this application.

[0285] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals.

[0286] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the terminal device 300. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0287] The wireless communication module 160 can provide solutions for wireless communication applications on the terminal device 300, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0288] The structure illustrated in Figure 14 does not constitute a specific limitation on the terminal device 300. The terminal device 300 may include more or fewer parts than illustrated, or combine some parts, or split some parts, or arrange different parts. The various parts illustrated may be implemented in hardware, software, or a combination of software and hardware.

[0289] Figure 15 illustrates a server 300 provided in an embodiment of this application. The server 300 may be a cloud server mentioned in the foregoing embodiments. As shown in Figure 15, the server 300 may include: a processor 210, a memory 220, an input / output device 230, a communication module 240, etc., and these components may be coupled via a bus.

[0290] Server 300 may have powerful computing resources, and its processor 210 may include one or more powerful processors, such as central processing unit (CPU), neural network processing unit (NPU), graphics processing unit (GPU), etc.

[0291] The processor 210 may include one or more interfaces, such as an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0292] The processor 210 may have a cache memory, which can be used to store instructions or data that the processor 210 has just used or that are used repeatedly. If the processor 210 needs to use the instruction or data again, it can directly retrieve it from the cache memory, which can reduce the waiting time of the processor 210 and improve the program running efficiency.

[0293] The processor 210 can also connect to external memory. This memory can be high-speed random access memory or non-volatile memory, such as a hard disk, flash memory, universal flash memory (UFS), etc. The memory can also be an external memory card, such as a Micro SD card.

[0294] The processor 210 is the computing core of the server 300, possessing powerful computing capabilities. Coupled with the memory 220, it can read and execute computer-readable instructions stored in the memory 220, running the operating system and various programs. Specifically, the CPU 210 can call programs stored in the memory 220, such as the implementation program of the edge-cloud collaborative computing power scheduling method provided in this application embodiment, and execute the instructions contained in that program.

[0295] The memory 220 may include high-speed random access memory, non-volatile memory, such as disk, flash memory, or other non-volatile solid-state storage devices. The memory 220 can be used to store various software programs and multiple sets of instructions. The memory 220 can store an operating system, such as Linux. The memory 220 can also store one or more programs, such as programs involved in patch creation, such as compilers and linkers. The memory 220 can also store the implementation program of the edge-cloud collaborative computing power scheduling method provided in the embodiments of this application.

[0296] Input / output device 230 may include devices such as a display screen, keyboard, and mouse, and can be used to receive user input and output program execution results to the user.

[0297] The communication module 240 may include a wired communication module and a wireless communication module. The wired communication module supports wired communication protocols such as Universal Serial Bus (USB), serial port, and Ethernet, communicating with other devices via physical communication cables. The wireless communication module may include 2G / 3G / 4G / 5G wireless communication modules, Wi-Fi communication modules, etc. The wireless communication module receives electromagnetic waves via an antenna, modulates and filters the electromagnetic wave signals, and sends the processed signal to the CPU 210. The wireless communication module can also receive signals to be transmitted from the CPU 210, modulate and amplify them, and then convert them into electromagnetic waves for radiation via the antenna.

[0298] The structure illustrated in Figure 15 does not constitute a limitation on server 300. Server 300 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The components illustrated may be implemented in hardware, software, or a combination of software and hardware.

[0299] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the above-described method embodiments.

[0300] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the human-computer interaction steps in the above-described method embodiments, such as receiving a question (query) through a voice assistant, outputting a reply text, displaying a resource service card, and playing / jumping to a page.

[0301] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the reasoning operation steps based on the text big model and multimodal big model in the above-described method embodiments, such as user intent recognition, resource search, and resource selection.

[0302] This application also provides a computer program product that, when executed by a processor, can implement all the steps in the above-described method embodiments.

[0303] This application also provides a computer program product that, when executed by a processor, can implement the human-computer interaction steps described in the above-described method embodiments.

[0304] This application also provides a computer program product that, when executed by a processor, can implement the reasoning operation steps based on the text large model and the multimodal large model in the above-described method embodiments.

[0305] This application also provides a chip system, which includes a processor coupled to a memory. The processor executes a computer program stored in the memory. When executed by the processor, the computer program can implement all the steps in the above-described method embodiments. The chip system can be a single chip or a chip module composed of multiple chips.

[0306] This application also provides a chip system, which includes a processor coupled to a memory. The processor executes a computer program stored in the memory. When executed by the processor, the computer program can implement the human-computer interaction steps described in the various method embodiments above. The chip system can be a single chip or a chip module composed of multiple chips.

[0307] This application also provides a chip system, which includes a processor coupled to a memory. The processor executes a computer program stored in the memory. When executed by the processor, the computer program can perform the inference operations based on the text-based large model and the multimodal large model described in the above-described method embodiments. The chip system can be a single chip or a chip module composed of multiple chips.

[0308] The term "user interface (UI)," or simply "interface," used in the specification and accompanying drawings of this application, refers to the medium through which an application or operating system interacts and exchanges information with the user. It facilitates the conversion between the internal form of information and a form acceptable to the user. The user interface of an application is written in source code using specific computer languages ​​such as Java or Extensible Markup Language (XML). This source code is parsed and rendered on the terminal device, ultimately presenting user-recognizable content, such as images, text, and buttons. Controls, also known as widgets, are the basic elements of the user interface. Typical controls include toolbars, menu bars, text boxes, buttons, scroll bars, images, and text. The attributes and content of controls in the interface are defined using tags or nodes, such as XML tags. <textview> 、 <imgview> 、 <videoview>Nodes define the controls contained in the interface. A node corresponds to a control or property in the interface, and after parsing and rendering, the node is presented as the content visible to the user. In addition, many applications, such as hybrid applications, often contain web pages within their interfaces. A web page, also known as a webpage, can be understood as a special control embedded in the application interface. Web pages are source code written in a specific computer language, such as Hypertext Markup Language (HTML), Cascading Style Sheets (CSS), JavaScript (JS), etc. Web page source code can be loaded and displayed as user-readable content by a browser or a web page display component with browser-like functionality. The specific content contained in a webpage is also defined through tags or nodes in the webpage source code; for example, HTML uses tags or nodes to define the content. 、 、 <video> 、 <canvas>Used to define the elements and attributes of a webpage.

[0309] The most common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of a terminal device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0310] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0311] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

[0312] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.< / canvas> < / video> < / videoview> < / imgview> < / textview>

Claims

1. A video resource search method, characterized in that, include: The terminal device receives the first dialogue content input by the first user, and the first dialogue content expresses the resource search intent in natural language. Based on the first dialogue content, the terminal device replies with a first resource fragment that matches the resource search intent; the first resource fragment is a video resource fragment, which is one or more episodes of a film or television work, or one or more film or television fragments from one episode of a film or television work. The first resource fragment whose response matches the resource search intent includes displaying a first response text and / or displaying a first service card; the first response text is used to describe the first resource fragment, and the first service card is used by the user to view the first resource fragment.

2. The method as described in claim 1, characterized in that, The method further includes: The terminal device sends the first dialogue content to the server; The terminal device receives indication information of the first resource segment from the server. The indication information of the first resource segment includes the identification information of the resource to which the first resource segment belongs, and the location information of the first resource segment in its resource. The location information includes the start time of the first resource segment.

3. The method as described in claim 1 or 2, characterized in that, The method further includes: the terminal device receiving the first reply text from the server.

4. The method according to any one of claims 1-3, characterized in that, The method further includes: The terminal device receives a first image from the server, the first image being generated based on the first resource fragment; The terminal device uses the first image as the cover image of the first service card.

5. The method as described in claim 4, characterized in that, The first image is the first frame of the first resource segment.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: the terminal device detecting a user operation applied to the first service card; jumping to the playback interface of the first resource segment and playing the first resource segment.

7. The method according to any one of claims 1-6, characterized in that, The rendered content on the first service card is used to describe the first resource fragment, including one or more of the following: the name, type, release time, creator, and cover image of the first resource fragment or its associated video resource.

8. The method according to any one of claims 1-7, characterized in that, The method further includes: The terminal device receives the first dialogue content input by a second user, who is different from the first user; Based on the first dialogue content input by the second user, the terminal device replies to the second user with a second resource fragment that matches the resource search intent; the second resource fragment is one or more episodes of a film or television work, or one or more film or television fragments from one episode of a film or television work; The second resource fragment whose response matches the resource search intent includes displaying a second response text and / or displaying a second service card; the second response text is used to describe the second resource fragment, the second service card is used for the user to watch the second resource fragment, the second resource fragment has different video content from the first resource fragment, or the second resource fragment has the same content as the first resource fragment but belongs to a different resource, or the second resource fragment has the same content as the first resource fragment but has a different clarity.

9. The method according to claim 8, characterized in that, The memory information or personal profile of the first user and the second user on the terminal device are different.

10. The method according to any one of claims 1-9, characterized in that, There are multiple first resource fragments that match the stated resource search intent; After the response matches the first resource fragment that matches the resource search intent, the method further includes: The terminal device receives a second dialogue content input by the first user, the second dialogue content expressing the intention to select a third resource fragment from a plurality of first resource fragments in natural language; According to the second dialogue content, the terminal device selects a third resource segment from a plurality of first resource segments, jumps to the playback interface of the third resource segment, and plays the third resource segment.

11. A video resource search method, characterized in that, include: The server receives the first dialogue content from the terminal device, and the first dialogue content expresses the resource search intent in natural language. The server determines a first resource fragment that matches the resource search intent based on the first dialogue content; the first resource fragment is a video resource fragment, which is one or more episodes of a film or television work, or one or more film or television fragments in one episode of a film or television work. The server returns indication information of the first resource fragment to the terminal device. The indication information includes the identification information of the resource to which the first resource fragment belongs, and the location information of the first resource fragment in its resource. The location information includes the start time of the first resource fragment.

12. The method as described in claim 11, characterized in that, Also includes: The server also returns the first response text to the terminal device, the first response text being used to describe the first resource fragment.

13. The method as described in claim 11 or 12, characterized in that, Also includes: The server also returns a first image to the terminal device, the first image being generated based on the first resource fragment.

14. The method as described in claim 13, characterized in that, The first image is the first frame of the first resource segment.

15. The method according to any one of claims 11-14, characterized in that, The server determines a first resource fragment that matches the resource search intent based on the first dialogue content, specifically including: The server uses the PGC search results and Internet search results associated with the first dialogue content to determine a first resource fragment that matches the resource search intent for the first dialogue content. The internet search results associated with the first dialogue content are obtained by searching the internet based on the first dialogue content. The professionally produced content (PGC) search results associated with the first dialogue content are obtained by searching the PGC resource library based on the first resource identifier. The first resource identifier is extracted from the internet search results. The PGC resource library records the resource fragments included in each different resource.

16. The method as described in claim 15, characterized in that, The server determines a first resource fragment that matches the resource search intent based on the first dialogue content, specifically including: The server first determines a third resource fragment that matches the resource search intent based on the Internet search results and the PGC search results; The server then uses a multimodal large model to process the third resource fragment, the internet search results, and the PGC search results to further locate the first resource fragment from the third resource fragment.

17. The method as described in claim 15 or 16, characterized in that, The server determines a first resource fragment that matches the resource search intent based on the first dialogue content, specifically including: The server analyzes the internet search results and the PGC search results based on the user profile of the first user to determine the first resource fragment.

18. The method according to any one of claims 15-17, characterized in that, Before the server determines a first resource fragment that matches the resource search intent based on the first dialogue content, the method further includes: The server uses the PGC search results to filter out inaccurate content from the internet search results.

19. The method according to any one of claims 11-18, characterized in that, There are multiple first resource fragments that match the resource search intent; after the server returns indication information of the first resource fragments to the terminal device, the method further includes: The server receives a second dialogue content from the terminal device, the second dialogue content expressing the intention to select resources in natural language; The server determines a third resource fragment that matches the resource selection intent from a plurality of first resource fragments based on the second dialogue content: The server returns indication information of the third resource fragment to the terminal device. The indication information of the third resource fragment is used to indicate the third resource fragment among a plurality of first resource fragments.

20. A terminal device, characterized in that, include: A processor, a memory, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of the method according to any one of claims 1-10.

21. A server, characterized in that, include: A processor, a memory, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of the method according to any one of claims 11-19.

22. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-19.

23. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-19.