A film and television search recommendation method, system, device and medium based on a large model

CN118467780BActive Publication Date: 2026-09-15E-SURFING DIGITAL LIFE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410574681.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-10
Publication Date
2026-09-15
Estimated Expiration
2044-05-10

AI Technical Summary

Technical Problem

相关技术中,通常依赖于基于关键词的搜索或元数据分类进行影视搜索,但是,在实际应用中发现,这些方法存在准确性和个性化程度方面的限制,导致搜索结果难以满足用户需求,影响了影视搜索推荐的效率

Benefits of technology

[0044]The embodiments of this application include at least the following beneficial effects: This application provides a method, system, device, and medium for film and television search and recommendation based on a large model. This solution automatically processes voice input data for speech recognition to obtain recognized text; performs intent recognition processing on the recognized text, and generates prompt words based on the intent recognition results combined with a multi-turn interaction algorithm to obtain search prompt words. This allows for a better understanding of user intent and preferences through intent recognition, thereby achieving a more intelligent search function. Furthermore, the combination with the multi-turn interaction algorithm enriches the search and recommendation experience. In addition, this solution inputs search prompt words into a large model for film and television resource search and recommendation processing to obtain search recommendation results; and displays and broadcasts the search recommendation results via voice. This allows for the use of a large model for film and television resource search and recommendation, improving search accuracy and providing personalized film and television content recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118467780B_ABST
    Figure CN118467780B_ABST
Patent Text Reader

Abstract

The application discloses a film and television search recommendation method, system, device and medium based on a large model. The method comprises the following steps: acquiring voice input data; performing automatic speech recognition processing on the voice input data to obtain recognized text; performing intent recognition processing on the recognized text, and generating a search prompt word according to an intent recognition result and a multi-round interaction algorithm; inputting the search prompt word into a large model to perform film and television resource search and recommendation processing, and obtaining a search recommendation result; and performing display and voice broadcast processing on the search recommendation result. The embodiment of the application can improve the accuracy of film and television search, provide personalized content recommendation, and can be widely applied to the field of artificial intelligence technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, system, device and medium for film and television search and recommendation based on a large model. Background Technology

[0002] With the rise of digitization of movies and TV series, the scale of film and television content libraries is growing rapidly. Related technologies typically rely on keyword-based search or metadata classification for film and television searches. However, in practical applications, these methods have been found to have limitations in accuracy and personalization, making it difficult for search results to meet user needs and impacting the efficiency of film and television search and recommendation.

[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention

[0004] The main objective of this application is to propose a film and television search and recommendation method, system, device, and medium based on a large model, which can improve the accuracy and efficiency of search and recommendation.

[0005] To achieve the above objectives, one aspect of this application proposes a film and television search and recommendation method based on a large model, the method comprising:

[0006] Acquire voice input data;

[0007] The voice input data is subjected to automatic speech recognition processing to obtain the recognized text;

[0008] The identified text is subjected to intent recognition processing, and prompt words are generated based on the intent recognition results and a multi-turn interaction algorithm to obtain search prompt words;

[0009] The search suggestions are input into a large model for film and television resource search and recommendation processing to obtain search and recommendation results.

[0010] The search recommendation results are displayed and read aloud via voice.

[0011] In some embodiments, the automatic speech recognition processing of the voice input data to obtain recognized text includes:

[0012] The voice input data is preprocessed to obtain preprocessed data;

[0013] The preprocessed data is subjected to feature extraction processing to obtain speech features;

[0014] The speech features are matched and decoded to obtain the recognized text.

[0015] In some embodiments, the process of performing intent recognition processing on the identified text and generating prompt words based on the intent recognition results combined with a multi-turn interaction algorithm to obtain search prompt words includes:

[0016] The identified text is subjected to intent recognition processing to obtain intent recognition results;

[0017] The recognized text is reconstructed based on the intent recognition result to obtain the reconstructed text.

[0018] Historical dialogue information is obtained based on a multi-round interaction algorithm;

[0019] The historical dialogue information and the reconstructed text are fused together to obtain the fused text;

[0020] Based on the fused text, prompt words are generated to obtain search prompt words.

[0021] In some embodiments, the step of generating search suggestion words based on the fused text includes:

[0022] Keywords are determined based on a prompt word template, and the keywords include roles, tasks, requirements, and prompts;

[0023] The fused text is combined with the keywords to generate search suggestion words.

[0024] In some embodiments, the step of inputting the search suggestions into a large model for film and television resource search and recommendation processing to obtain search recommendation results includes:

[0025] Access to film and television resource library;

[0026] The search suggestions are input into a large model, and the search suggestions and the film and television resource library are processed by vector representation and matching to output the search results.

[0027] Obtain historical behavior data;

[0028] The search results are processed using multi-dimensional recommendation methods based on the historical behavior data to obtain the search recommendation results.

[0029] In some embodiments, the process of displaying and voice-reading the search recommendation results includes:

[0030] The search recommendation results are formatted, parsed, and extracted to obtain interactive text and film and television resources;

[0031] The video resources are displayed and the interactive text is read aloud.

[0032] In some embodiments, the voice broadcasting processing of the interactive text includes:

[0033] The interactive text is segmented and tagged to obtain tagged text;

[0034] The marked text is processed by phoneme mapping to obtain text phonemes;

[0035] The text phonemes are processed by sound synthesis by selecting sound data or synthesis rules to obtain audio data, which is then transmitted and played.

[0036] To achieve the above objectives, another aspect of this application proposes a film and television search and recommendation system based on a large model, the system comprising:

[0037] The first module is used to acquire voice input data;

[0038] The second module is used to perform automatic speech recognition processing on the voice input data to obtain the recognized text;

[0039] The third module is used to perform intent recognition processing on the identified text, and to generate prompt words based on the intent recognition results and a multi-round interaction algorithm to obtain search prompt words.

[0040] The fourth module is used to input the search suggestions into the large model for film and television resource search and recommendation processing, and to obtain search recommendation results;

[0041] The fifth module is used to display and voice-broadcast the search recommendation results.

[0042] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.

[0043] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.

[0044] The embodiments of this application include at least the following beneficial effects: This application provides a method, system, device, and medium for film and television search and recommendation based on a large model. This solution automatically processes voice input data for speech recognition to obtain recognized text; performs intent recognition processing on the recognized text, and generates prompt words based on the intent recognition results combined with a multi-turn interaction algorithm to obtain search prompt words. This allows for a better understanding of user intent and preferences through intent recognition, thereby achieving a more intelligent search function. Furthermore, the combination with the multi-turn interaction algorithm enriches the search and recommendation experience. In addition, this solution inputs search prompt words into a large model for film and television resource search and recommendation processing to obtain search recommendation results; and displays and broadcasts the search recommendation results via voice. This allows for the use of a large model for film and television resource search and recommendation, improving search accuracy and providing personalized film and television content recommendations. Attached Figure Description

[0045] Figure 1 This is a flowchart of a film and television search and recommendation method based on a large model provided in an embodiment of this application;

[0046] Figure 2 yes Figure 1 The flowchart of step S102 in the document;

[0047] Figure 3 yes Figure 1 The flowchart of step S103 in the process;

[0048] Figure 4 yes Figure 3 The flowchart of step S305 in the text;

[0049] Figure 5 yes Figure 1 The flowchart of step S104 in the process;

[0050] Figure 6 yes Figure 1 The flowchart of step S105 in the process;

[0051] Figure 7 yes Figure 6 The flowchart of step S602 in the document;

[0052] Figure 8 This is a flowchart illustrating a specific implementation of a film and television search and recommendation method based on a large model, as provided in an embodiment of this application.

[0053] Figure 9 This is a schematic diagram of the structure of a film and television search and recommendation system based on a large model provided in an embodiment of this application;

[0054] Figure 10 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of systems and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0056] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if” or “when” as used herein may be interpreted as “when…” or “in response to determination.”

[0057] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0059] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained first. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.

[0060] Artificial Intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0061] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software primarily includes computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0062] Large models, also known as foundation models, are machine learning models with a large number of parameters and complex structures. They are capable of processing massive amounts of data and performing various complex tasks, such as natural language processing, computer vision, and speech recognition. The purpose of large models is to improve their expressive power and predictive performance, enabling them to handle more complex tasks and data.

[0063] Automatic Speech Recognition (ASR), also known as automated speech recognition, aims to convert the lexical content of human speech into computer-readable input, such as keystrokes, binary codes, or character sequences. This differs from speaker recognition and speaker verification, which attempt to identify or verify the speaker rather than the lexical content of the speech.

[0064] Text-to-speech (TTS) technology is a technology that converts computer-generated or externally input text information into fluent, understandable spoken Chinese (or other languages) output. It belongs to the field of speech synthesis technology.

[0065] In related technologies, methods for searching film and television content typically rely on keyword-based search or metadata classification; however, these methods have limitations in terms of accuracy and personalization. With the massive expansion of digital entertainment content, existing keyword-based search and metadata classification methods are no longer sufficient to meet user needs. Therefore, this application aims to leverage the potential of large-scale pre-trained models to provide more intelligent film and television search and recommendation services to meet users' ever-growing demand for movie and television content.

[0066] In view of this, this application provides a method, system, device, and medium for film and television search and recommendation based on a large model. This solution automatically processes voice input data through speech recognition to obtain recognized text; it then processes the recognized text through intent recognition and generates prompt words based on the intent recognition results and a multi-turn interaction algorithm, resulting in search prompt words. This allows for a better understanding of user intent and preferences through intent recognition, thereby achieving a more intelligent search function. Furthermore, the multi-turn interaction algorithm enriches the search and recommendation experience. Additionally, this solution inputs the search prompt words into a large model for film and television resource search and recommendation processing to obtain search and recommendation results. The search and recommendation results are then displayed and broadcast via voice. This approach utilizes a large model for film and television resource search and recommendation, improving search accuracy and providing personalized film and television content recommendations.

[0067] This application provides a film and television search and recommendation method based on a large model, relating to the field of artificial intelligence technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited thereto. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing a film and television search and recommendation method based on a large model, but is not limited to the above forms.

[0068] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0069] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0070] Figure 1 This is an optional flowchart of a film and television search and recommendation method based on a large model provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S101 to S105.

[0071] Step S101: Obtain voice input data;

[0072] Step S102: Perform automatic speech recognition processing on the voice input data to obtain the recognized text;

[0073] Step S103: Perform intent recognition processing on the identified text, and generate prompt words based on the intent recognition results and a multi-round interaction algorithm to obtain search prompt words;

[0074] Step S104: Input the search suggestions into the large model for film and television resource search and recommendation processing to obtain search recommendation results;

[0075] Step S105: Display and voice broadcast the search recommendation results.

[0076] Steps S101 to S105 of this embodiment involve acquiring user voice input data, converting the voice input data into text format using automatic speech recognition technology to obtain recognized text, and then transmitting the recognized text to a large-scale model for intent recognition processing. Based on the intent recognition results, a multi-turn interaction algorithm is used to generate prompt words, which can more accurately describe the user's intent and generate search prompt words suitable for different scenarios. The finally generated search prompt words are then input into the large-scale model for film and television resource search and recommendation processing to obtain search recommendation results. These results include film and television titles and interactive text. This embodiment uses the large-scale model to search for film and television resources and interactive text related to the user's intent, displays the found film and television resources, and broadcasts the interactive text via voice. This embodiment, by utilizing a large-scale model for film and television resource search and recommendation, can provide more intelligent film and television search and recommendation services to meet users' ever-growing demand for movie and television content.

[0077] In step S101 of some embodiments, the user's voice input data can be acquired through the microphone of the smart terminal. This voice input data can be in the form of a question. For example, when a user is using a smart robot or smart software program capable of human-computer interaction, the user asks, "Can you recommend a recent romance movie?" This embodiment of the application acquires this voice input data, processes it through automatic speech recognition and intent recognition, and then inputs it into a large model for film and television resource search and recommendation. The search and recommendation results that meet the user's needs are then displayed and voice-read aloud, providing a more intelligent film and television search and recommendation service through the large model.

[0078] Please see Figure 2 In some embodiments, step S102 may include, but is not limited to, steps S201 to S203:

[0079] Step S201: Preprocess the voice input data to obtain preprocessed data;

[0080] Step S202: Perform feature extraction processing on the preprocessed data to obtain speech features;

[0081] Step S203: Perform matching and decoding processing on the speech features to obtain the recognized text.

[0082] In step S201 of some embodiments, the speech input data is preprocessed to obtain preprocessed data. The preprocessing includes noise and interference removal to reduce background noise and interference in the speech input data. Specifically, a digital filter can be used to filter the speech input data to remove noise within a specific frequency range. Alternatively, the spectrum of the speech input data can be analyzed, the noise spectrum can be compared with the speech spectrum, and then the noise spectrum can be subtracted to recover clear speech input data. Alternatively, a speech denoising model can be trained using machine learning techniques. By using training data with pairs of noisy and clean speech, a model can be trained to learn the relationship between noise and speech and to remove noise.

[0083] In step S202 of some embodiments, the preprocessed data is processed by a speech recognition module to extract speech-related features and obtain speech features. These speech features include Mel-frequency cepstral coefficients, which can be extracted by the speech recognition module through processes such as fast Fourier transform, filter bank, logarithmic operation, discrete cosine transform, and dynamic feature extraction.

[0084] In step S203 of some embodiments, the feature vectors are matched and decoded using a speech model and an acoustic model to obtain the recognized text. The speech model is a knowledge representation of a sequence of characters, and the acoustic model is a knowledge representation of differences in acoustic, phonetic, and environmental variables. In this embodiment, the input speech feature vector sequence is matched and decoded using a speech model and an acoustic model, thereby converting it into a character sequence to obtain the recognized text.

[0085] In this embodiment, preprocessed data is obtained by preprocessing the voice input data, and voice features are obtained by feature extraction of the preprocessed data. Finally, the voice features are matched and decoded to obtain the recognized text. This enables automatic voice recognition processing of voice input data, providing a data foundation for subsequent intent recognition and prompt word generation.

[0086] Please see Figure 3 In some embodiments, step S103 may include, but is not limited to, steps S301 to S305:

[0087] Step S301: Perform intent recognition processing on the identified text to obtain intent recognition results;

[0088] Step S302: Reconstruct the recognized text based on the intent recognition result to obtain the reconstructed text;

[0089] Step S303: Obtain historical dialogue information based on a multi-round interaction algorithm;

[0090] Step S304: The historical dialogue information and the reconstructed text are fused together to obtain fused text;

[0091] Step S305: Perform prompt word generation processing based on the fused text to obtain search prompt words.

[0092] In step S301 of some embodiments, intent recognition results are obtained by performing intent recognition processing on the identified text; wherein, the identified text can be transmitted to a large model for intent recognition, and specific entities and intents can be identified based on a predefined set of rules. The large model is trained to learn patterns for extracting intents and entities from the text and can make inferences to process new user inputs.

[0093] In step S302 of some embodiments, the identified text is reconstructed based on the intent recognition result to obtain reconstructed text. The identified text can be reconstructed based on the intent recognition result to more accurately describe the user's intent.

[0094] In step S303 of some embodiments, historical dialogue information is obtained according to a multi-turn interaction algorithm. In this embodiment of the application, human-computer interaction can be performed through a human-computer interaction interface. The human-computer interaction can include text input and voice input. Through a multi-turn dialogue scenario, historical dialogue records of multi-turn conversations are obtained.

[0095] In step S304 of some embodiments, historical dialogue information and reconstructed text are fused to obtain fused text. Human-computer interaction technology can be used to further expand the recognized text through multi-turn interactions. For example, after a user inputs "Can you recommend a recent romance movie?", a multi-turn interaction algorithm can be used to engage in multiple rounds of dialogue with the user, further narrowing down the specific year range and inquiring about the movie's style or theme, actors or directors, and place of origin. This, combined with historical dialogue information, allows for fusion processing of the reconstructed text, resulting in the fused text. This application embodiment allows users to interact more deeply with the system through multi-turn interaction algorithms, providing a richer search and recommendation experience and helping to meet users' detailed needs, especially when dealing with more complex queries, thereby improving the efficiency of film and television search. It should be noted that when multi-turn dialogue is not performed, the reconstructed text can be directly combined with a prompt word generation template to generate prompt words suitable for different scenarios.

[0096] Please see Figure 4 In some embodiments, step S305 may include, but is not limited to, steps S401 to S402:

[0097] Step S401: Determine keywords based on the prompt word template, wherein the keywords include role, task, requirement and prompt;

[0098] Step S402: Combine the fused text with the keywords to generate search suggestion words.

[0099] In step S401 of some embodiments, a prompt word generation template is pre-set, thereby generating prompt words suitable for different scenarios based on the prompt word generation template. Corresponding keywords are set for the large model, including roles, tasks, requirements, and prompts. Roles act as prior knowledge, providing necessary background information for the large model; tasks clarify the specific tasks the large model needs to complete and the limitations of the generated content; requirements constrain the final format of the generated large model; and prompts are instances of the final generated format, used to guide the generation process of the large model. Finally, the fused text is used as the user's question, combined with the keywords set above, for generation processing to obtain search prompt words. For example, a specific form of a search prompt word is shown below:

[0100] "Role: Imagine you are an experienced film expert."

[0101] Task: Please generate a list of videos related to the user's question and generate a voice narration.

[0102] Requirements: Return the audio narration (Title) and video list (Program_Names) in JSON format.

[0103] Hint: ```json{{"T it le":"","Progrom_Names":[]}}```

[0104] User question: Can you recommend a romance movie from recent years?

[0105] In the application embodiment, by setting corresponding keywords for the large model, search suggestion words suitable for different scenarios are generated by combining fused text and keywords, thereby providing a data foundation for subsequent film and television search recommendation processing of the large model.

[0106] Please see Figure 5 In some embodiments, step S104 may also include, but is not limited to, steps S501 to S504:

[0107] Step S501: Obtain the film and television resource library;

[0108] Step S502: Input the search suggestion words into the large model, perform vector representation and matching processing on the search suggestion words and the film and television resource library, and output the search results;

[0109] Step S503: Obtain historical behavior data;

[0110] Step S504: Perform multi-dimensional recommendation processing on the search results based on the historical behavior data to obtain search recommendation results.

[0111] In step S501 of some embodiments, the corresponding film and television resource library can be obtained through external interfaces and other technologies. The film and television resource library includes content such as movies and television programs.

[0112] In step S502 of some embodiments, the search suggestion words are input into a large model, and the large model converts the search suggestion words and the film and television resource library into continuous vector representations for matching and recommendation. Then, the corresponding film and television resources are obtained by matching algorithms such as vector similarity calculation, and the search results are output.

[0113] In step S503 of some embodiments, historical behavior data is obtained, which includes the user's historical search and viewing behavior data. The user's interests and preferences can be determined through the historical behavior data, and personalized recommendations can be provided based on this information.

[0114] In step S504 of some embodiments, the search results are processed for multi-dimensional recommendation based on historical behavior data. Multiple dimensions can be considered at the same time, such as content type, actors, directors, user ratings, etc. For example, search results with a score of eight or higher are recommended based on user ratings to provide more accurate recommendation results.

[0115] This application embodiment uses a large model to search for film and television resources, which can match user queries and vector representations of content, and recommend search results through multi-dimensional recommendation technology, providing personalized and intelligent search and recommendation functions, and improving the efficiency of film and television search and recommendation.

[0116] Please see Figure 6 In some embodiments, step S105 is not limited to including steps S601 to S602:

[0117] Step S601: The search recommendation results are formatted and parsed to extract interactive text and film and television resources;

[0118] Step S602: Display the film and television resources and perform voice broadcasting on the interactive text.

[0119] In step S601 of some embodiments, a formatting operation is performed on the search results to generate standardized JSON output, which includes interactive text and movie / TV show titles. Then, the movie / TV show titles are extracted by parsing the JSON text in order to search for relevant movie / TV show resources. The movie / TV show titles are movie / TV show titles recommended by a large model, and the interactive text is feedback text for human-computer interaction, such as interactive content such as "The following are recommended movies / TV shows for you", which can be read aloud by voice or displayed in text form on the human-computer interaction interface.

[0120] In step S602 of some embodiments, film and television resources can be displayed and processed through human-computer interaction interfaces, displays, projections, etc., and text-to-speech technology can be used to process the interactive text into voice broadcasts. This application embodiment provides users with personalized recommended content and improves the efficiency of film and television search and recommendation by displaying and broadcasting search and recommendation results into voice broadcasts.

[0121] Please see Figure 7 In some embodiments, step S602 may include, but is not limited to, steps S701 to S703:

[0122] Step S701: Perform word segmentation and tagging on the interactive text to obtain tagged text;

[0123] Step S702: Perform phoneme mapping processing on the marked text to obtain text phonemes;

[0124] Step S703: The text phonemes are processed by sound synthesis by selecting sound data or synthesis rules to obtain audio data and transmit it for playback.

[0125] In step S701 of some embodiments, the input interactive text is processed, including word segmentation, grammatical analysis, and speech tokenization. By performing word segmentation and tokenization on the interactive text, it is helpful to determine the phonetic morphemes, grammatical structures, and pronunciation rules in the text, providing a foundation for subsequent speech recognition.

[0126] In step S702 of some embodiments, the marked text is processed by phoneme mapping, which maps words, phrases and sentences in the text to speech units to obtain text phonemes. Here, a phoneme is a basic building block of speech, and each phoneme usually corresponds to a specific sound.

[0127] In step S703 of some embodiments, the text phonemes are processed by audio synthesis by selecting appropriate sound data or synthesis rules. This may include selecting suitable speech segments or sound units to construct the pronunciation of the phonemes, and finally performing audio synthesis, including splicing the selected phonemes and applying sound effects, such as adjusting pitch, volume, and speech rate, to make the sound sound natural. This yields audio data for transmission and playback. This audio data is typically a digital audio stream or audio file, which can be stored as a file, transmitted to a player, or transmitted to the user in real time. Generally, the audio data is represented digitally and uses standard audio formats such as WAV, MP3, or OGG.

[0128] The following section provides a detailed description and explanation of the solutions in the embodiments of this application, taking into account specific application scenarios:

[0129] This application's embodiments can be applied to scenarios such as streaming media platforms, social media, video sharing websites, and advertising recommendation systems. Streaming media platforms are crucial providers of film and television content; the methods in this application's embodiments can be applied to the search and recommendation systems of streaming media platforms to improve user experience. Social media is an important platform for users to share and exchange film and television content; the methods in this application's embodiments can be applied to the search and recommendation systems of social media to provide users with more accurate and personalized content recommendations. Video sharing websites are important platforms for users to watch and share film and television content; the methods in this application's embodiments can be applied to the search and recommendation systems of video sharing websites to improve the efficiency of users finding and watching film and television content. Advertising recommendation systems need to understand users' interests and behaviors in order to accurately push advertisements; the methods in this application's embodiments can be applied to advertising recommendation systems to improve advertisement click-through rates and conversion rates.

[0130] Reference Figure 8This system acquires a user-inputted voice question, performs automatic speech recognition on the question, converts the speech into text, and then inputs the text into a large-scale model for intent recognition and rewrites the user's question. Human-computer interaction technology is used to determine if a multi-turn dialogue occurred. If so, historical dialogue information is retrieved, and combined with the reconstructed question and historical dialogue information, prompts suitable for different scenarios are generated based on a prompt word template. If no multi-turn dialogue occurred, prompts are directly generated based on the reconstructed question and the prompt word template. Finally, the generated prompts are input into a large-scale model for film and television search and recommendation. The output is formatted as JSON and parsed to obtain the interactive text and film / television titles. The interactive text is then converted into audio, and the corresponding film / television resources are searched based on the film / television titles. These resources are displayed, and the interactive text is read aloud, completing the film and television resource search and recommendation process. This embodiment of the application can improve the efficiency, accuracy, and user satisfaction of film and television search and recommendation systems, providing users with a more intelligent, personalized, and entertaining film and television content experience.

[0131] Please see Figure 9 This application also provides a film and television search and recommendation system based on a large model, which can implement the above-mentioned film and television search and recommendation method based on a large model. The system includes:

[0132] The first module 901 is used to acquire voice input data;

[0133] The second module 902 is used to perform automatic speech recognition processing on the voice input data to obtain the recognized text.

[0134] The third module 903 is used to perform intent recognition processing on the identified text, and to perform prompt word generation processing based on the intent recognition result and a multi-round interaction algorithm to obtain search prompt words;

[0135] The fourth module 904 is used to input the search suggestions into the large model for film and television resource search and recommendation processing, and to obtain search recommendation results;

[0136] The fifth module 905 is used to display and voice broadcast the search recommendation results.

[0137] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0138] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned large-model-based film and television search and recommendation method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0139] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0140] Please see Figure 10 , Figure 10 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0141] The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0142] The memory 1002 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1002 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001 to implement a large-model-based film and television search recommendation method according to an embodiment of this application.

[0143] Input / output interface 1003 is used to implement information input and output;

[0144] The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0145] Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004);

[0146] The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.

[0147] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned large-model-based film and television search and recommendation method.

[0148] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0149] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0150] This application provides a method, system, device, and medium for film and television search and recommendation based on a large model. This solution automatically processes voice input data into recognized text; performs intent recognition processing on the recognized text; and generates prompt words based on the intent recognition results combined with a multi-turn interaction algorithm to obtain search prompt words. Intent recognition allows for a better understanding of user intent and preferences, resulting in a more intelligent search function. Furthermore, the multi-turn interaction algorithm enriches the search and recommendation experience. Additionally, this solution inputs the search prompt words into a large model for film and television resource search and recommendation, obtaining search and recommendation results. The search and recommendation results are then displayed and broadcast via voice. This leverages the large model for film and television resource search and recommendation, improving search accuracy and providing personalized film and television content recommendations. This helps the film and television search field continuously adapt to the development of digital entertainment content and meet the ever-changing needs of users.

[0151] This application provides a film and television search and recommendation method based on a large-scale model. This method leverages the massive knowledge, intent recognition, and multi-turn dialogue capabilities of the large-scale model to identify user intent and recommend suitable film and television resources, thereby enhancing the user's interactive experience. This application uses a large-scale model to search for movies and television programs and matches user queries with vector representations of content, providing personalized and intelligent search and recommendation functions. This application can also utilize users' historical search and viewing behavior data to determine user interests and preferences, and provide personalized recommendations based on this information. This application also provides a user-friendly multi-turn dialogue interface to support users in progressively refining their queries and providing feedback, thereby improving search results. This application can also infer user query intent and preferences, achieving more intelligent search and recommendation by transforming user queries and content descriptions into continuous vector representations for matching and recommendation, and simultaneously considering multiple dimensions for personalized recommendations to provide more accurate recommendation results.

[0152] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0153] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0154] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0155] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0156] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0157] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0158] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0159] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0160] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0161] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0162] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A film and television search and recommendation method based on a large model, characterized in that, The method includes: Acquire voice input data; The voice input data is subjected to automatic speech recognition processing to obtain the recognized text; The identified text is processed for intent recognition, and prompt words are generated based on the intent recognition results and a multi-turn interaction algorithm to obtain search prompt words. The search suggestions are input into a large model for film and television resource search and recommendation processing to obtain search and recommendation results. The search recommendation results are displayed and read aloud via voice. The process of performing intent recognition processing on the identified text, and generating prompt words based on the intent recognition results and a multi-turn interaction algorithm, yields search prompt words, including: The identified text is subjected to intent recognition processing to obtain intent recognition results; The recognized text is reconstructed based on the intent recognition result to obtain the reconstructed text; Historical dialogue information is obtained based on a multi-round interaction algorithm; The historical dialogue information and the reconstructed text are fused together to obtain the fused text; Based on the fused text, prompt words are generated to obtain search prompt words; The step of generating search suggestion words based on the fused text includes: Keywords are determined based on the prompt word template. The keywords include roles, tasks, requirements, and prompts. The roles serve as prior knowledge, providing necessary background information for the large model. The tasks clarify the specific tasks that the large model needs to complete and the limitations of the generated content. The requirements constrain the final format of the large model. The prompts are instances of the final generated format, used to guide the generation process of the large model. The fused text is combined with the keywords to generate search suggestion words.

2. The method according to claim 1, characterized in that, The automatic speech recognition processing of the voice input data to obtain the recognized text includes: The voice input data is preprocessed to obtain preprocessed data; The preprocessed data is subjected to feature extraction processing to obtain speech features; The speech features are matched and decoded to obtain the recognized text.

3. The method according to claim 1, characterized in that, The step of inputting the search suggestions into a large model for film and television resource search and recommendation processing to obtain search recommendation results includes: Access to film and television resource library; The search suggestions are input into a large model, and the search suggestions and the film and television resource library are processed by vector representation and matching to output the search results. Obtain historical behavior data; The search results are processed using multi-dimensional recommendation methods based on the historical behavior data to obtain the search recommendation results.

4. The method according to any one of claims 1 to 3, characterized in that, The process of displaying and voice-reading the search recommendation results includes: The search recommendation results are formatted, parsed, and extracted to obtain interactive text and film and television resources; The video resources are displayed and the interactive text is read aloud.

5. The method according to claim 4, characterized in that, The process of performing voice broadcasting on the interactive text includes: The interactive text is segmented and tagged to obtain tagged text; The marked text is processed by phoneme mapping to obtain text phonemes; The text phonemes are processed by sound synthesis by selecting sound data or synthesis rules to obtain audio data, which is then transmitted and played.

6. A film and television search and recommendation system based on a large model, characterized in that, The large-model-based film and television search and recommendation system is applied to the large-model-based film and television search and recommendation method as described in claim 1, wherein the system comprises: The first module is used to acquire voice input data; The second module is used to perform automatic speech recognition processing on the voice input data to obtain the recognized text; The third module is used to perform intent recognition processing on the identified text, and to generate prompt words based on the intent recognition results and a multi-round interaction algorithm to obtain search prompt words. The fourth module is used to input the search suggestions into the large model for film and television resource search and recommendation processing, and to obtain search recommendation results; The fifth module is used to display and voice-broadcast the search recommendation results.

7. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Search result broadcasting method and device based on artificial intelligence

    CN105653738A

  • Recommendation model training method and device, search text recommendation method and device and storage medium

    CN110727785A

  • Man-machine interaction and content search method and device, equipment and storage medium

    CN110765312A

  • Drawing retrieval method and device, electronic equipment and storage medium

    CN117972132A