Display device, method and related device

By collecting and analyzing user voice data through display devices, personalized travel recommendations are generated, solving the problems of inconvenient information collection and poor interactive experience in multi-person travel planning, and realizing intelligent travel planning driven by natural dialogue.

CN121979476APending Publication Date: 2026-05-05SHENZHEN TAILIWEI INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN TAILIWEI INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-01-27
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In the process of multi-person travel planning, existing technologies suffer from inconvenient information collection, poor interactive experience, homogenized recommendation results, and low efficiency in multi-person collaborative discussions.

Method used

By collecting users' voice conversations through display devices, performing preliminary recognition and intent analysis, generating personalized travel recommendations, and displaying them on a large screen to enable collaborative decision-making among multiple users.

Benefits of technology

It improves the efficiency and accuracy of travel recommendations, can automatically process travel needs in natural family conversations, generate personalized plans, and improve the efficiency of multi-person collaborative decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121979476A_ABST
    Figure CN121979476A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a display device and method and a related device, and relates to the technical field of display, and the method comprises the steps: carrying out the preliminary recognition of the voice, collected by a sound collection device, of user communication in a scene of multi-person discussion; when it is determined that the identification result comprises the tourism related information, obtaining tourism intentions of all the users; generating tourism recommendation information based on the tourism intention of each user; and controlling a display screen to display the tourism recommendation information. According to the mode, the tourism demand of the user can be recognized based on the communication voice of the user, and the tourism scheme considering the individual demand of the family member can be generated through the recommendation algorithm. Finally, multi-person collaborative decision-making is realized through large-screen display, and the efficiency and accuracy of tourism planning are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of displays, and more particularly to a display device, method and related apparatus. Background Technology

[0002] When planning a trip, users typically input keywords via their mobile phones. The system then matches and recommends travel itineraries based on these keywords and displays the results on the user's phone. However, when multiple people need to view these recommendations, they often have to switch phones or share them on social media, leading to inefficiency in collaborative discussions and a poor user experience. Summary of the Invention

[0003] This application provides a display device, method, and related apparatus to improve the efficiency and accuracy of travel recommendations.

[0004] In a first aspect, embodiments of this application provide a display device, the display device comprising:

[0005] The display screen is configured to display images.

[0006] A sound acquisition device, configured to acquire sound;

[0007] A controller connected to the display screen is configured to:

[0008] In scenarios involving multiple people discussing, the voice recording device is used to perform preliminary recognition of the users' conversations.

[0009] When the identification results include tourism-related information, obtain the tourism intent of each user;

[0010] Based on the travel intentions of each user, travel recommendation information is generated;

[0011] Control the display screen to show the travel recommendation information.

[0012] In some embodiments, the controller is configured to:

[0013] The speech is divided into multiple speech segments; each speech segment includes only the speech of one user.

[0014] The multiple speech segments are clustered to obtain the clustered segments corresponding to each user, and the clustered segments are sorted according to time order to obtain the target speech segments corresponding to each user.

[0015] The travel intentions of each user are obtained by performing intent recognition on the target speech segments of each user.

[0016] In some embodiments, the controller is configured to:

[0017] The target speech segment is subjected to speech recognition to obtain the text corresponding to the target speech segment;

[0018] The text is subjected to intent recognition to obtain the first travel intent;

[0019] Intent reasoning is performed on the text based on a tourism knowledge graph to obtain the implicit second tourism intent of the text.

[0020] The first travel intention and the second travel intention are merged to obtain the travel intention.

[0021] In some embodiments, the controller is configured to:

[0022] Obtain the entities related to tourism from the text;

[0023] The entity is mapped to the corresponding node in the tourism knowledge graph, and the second tourism intention is obtained by reasoning based on the user's historical tourism information starting from the node.

[0024] In some embodiments, the controller is configured to:

[0025] Align the first travel intention and the second travel intention;

[0026] If the first travel intention and the second travel intention match each other, then the second travel intention is supplemented based on the second travel intention to obtain the travel intention;

[0027] If the first travel intention and the second travel intention conflict, the first travel intention is modified based on the second travel intention to obtain the second travel intention.

[0028] In some embodiments, the controller is configured to:

[0029] Obtain the voiceprint features of each user, and determine the identity of each user based on the voiceprint features;

[0030] Based on the identity identifier, determine the weight of each user's corresponding travel intention;

[0031] Based on the weights, travel intentions, and identity identifiers, the travel propositions corresponding to each user are determined.

[0032] A tourism recommendation network is constructed based on the tourism propositions of each user, and a group tourism consensus vector for each user is determined based on the tourism recommendation network.

[0033] The tourism recommendation information is generated based on the group tourism consensus vector.

[0034] In some embodiments, the controller is configured to:

[0035] Determine the degree of controversy for each dimension in the aforementioned group tourism consensus vector;

[0036] When the degree of controversy is less than or equal to a preset value, a first recommended option is generated based on the corresponding dimensional consensus.

[0037] When the degree of controversy exceeds the preset value, multiple second recommendation options are generated based on the corresponding dimension of controversy.

[0038] The first recommended option and multiple second recommended options are combined to obtain the travel recommendation information.

[0039] In some embodiments, the controller is configured to:

[0040] Get the current environment status and displayed content;

[0041] Based on the environmental conditions and the displayed content, determine the target time for displaying the tourism recommendation information;

[0042] At the target time, the display screen is controlled to show the travel recommendation information.

[0043] In some embodiments, the controller is configured to:

[0044] Obtain user feedback instructions regarding the travel recommendation information;

[0045] The travel recommendation information is adjusted based on the feedback instructions.

[0046] Secondly, embodiments of this application provide a display method, including:

[0047] In scenarios involving multiple people in a discussion, the system collects the voice recordings of users' conversations and performs content recognition on the voice recordings.

[0048] When the content recognition results are determined to include tourism-related information, the tourism intent of each user is obtained.

[0049] Based on the travel intentions of each user, travel recommendation information is generated;

[0050] Control the display screen to show the travel recommendation information.

[0051] Thirdly, this application provides an electronic device, including: a memory and a processor;

[0052] The memory is used to store computer instructions; the processor is used to execute the computer instructions stored in the memory to implement the method of any of the second aspects.

[0053] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method of any of the second aspects.

[0054] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method of any one of the second aspects.

[0055] The display device, method, and related apparatus provided in this application, in a scenario of multi-person discussion, perform preliminary recognition of the voice collected by the sound acquisition device; when it is determined that the recognition result includes tourism-related information, the travel intentions of each user are obtained; based on the travel intentions of each user, tourism recommendation information is generated; and the display screen is controlled to display the tourism recommendation information. This method can recognize users' tourism needs based on their voice conversations, and through recommendation algorithms, can generate tourism plans that take into account the personalized needs of family members. Finally, multi-person collaborative decision-making is achieved through large-screen display, significantly improving the efficiency and accuracy of tourism planning. It effectively solves the problems of inconvenient information collection, poor interactive experience, and homogenized recommendation results in existing technologies, realizing a new paradigm of intelligent tourism planning driven by natural dialogue. Attached Figure Description

[0056] Figure 1 This is a schematic diagram of a display scenario provided in an embodiment of this application;

[0057] Figure 2 This is a schematic diagram of the structure of a display device provided in an embodiment of this application;

[0058] Figure 3 A flowchart illustrating a display method provided in an embodiment of this application. Figure 1 ;

[0059] Figure 4 A flowchart illustrating a display method provided in an embodiment of this application. Figure 2 ;

[0060] Figure 5 This is a schematic diagram of the structure of a display device provided in an embodiment of this application;

[0061] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0063] In the embodiments of this application, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect, without limiting their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" do not necessarily imply that they are different.

[0064] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0065] In modern family travel planning scenarios, family members typically discuss their travel intentions through everyday conversations. For example, during a family dinner, parents might mention, "I've been under a lot of work pressure lately, and I'd like to take the kids to the beach to relax," while the children might add, "I hope there's a water park and childcare services nearby," and grandparents might focus on "whether there are adequate medical facilities nearby."

[0066] Current tourism planning solutions mainly rely on the following technical paths: (1) Users manually input their tourism needs through mobile applications or web pages, including parameters such as destination, budget, and time; (2) The system recommends tourism solutions based on keyword matching, and the recommendation logic mainly relies on attraction tags and basic filtering conditions; (3) The recommendation results are displayed on a small mobile phone screen, and multiple people need to switch devices or forward them through social software when viewing them.

[0067] However, in the above solutions, the collection of tourism demand relies on users' active input, which cannot capture the implicit needs in the conversation; furthermore, the display of recommendation results is limited by the small screen of mobile devices, resulting in low efficiency for multi-person collaborative discussions.

[0068] To address the aforementioned issues, embodiments of this application provide a display device, method, and related apparatus that integrate voice acquisition, natural language processing, and personalized recommendation technologies to achieve fully automated processing of the entire process of automatically extracting travel needs from natural family conversations and generating personalized solutions, effectively improving the efficiency and accuracy of travel recommendations.

[0069] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.

[0070] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0071] Figure 1 A scenario illustration provided for this application, such as Figure 1 As shown, in scenarios involving multiple users communicating, the display device can collect the voices of users during their conversations and analyze the collected voices.

[0072] When the display device confirms that each user has a travel need during the communication process, it can summarize and analyze the travel preferences of each user, generate travel recommendation plans, and display the recommended travel plans on the display screen.

[0073] Figure 2 This is a schematic diagram of the structure of a display device 20 provided in an embodiment of this application, as shown below. Figure 2 As shown, it includes: a display 21, a voice acquisition device 22, and a controller 23.

[0074] The display 22 is configured to display a screen.

[0075] Voice acquisition device 22 is configured to acquire user voice.

[0076] The controller 23 is configured to generate a travel recommendation plan based on the collected user voice, convert the travel recommendation plan into a corresponding video signal, and send the video signal to the display 21 so that the display 21 can display the image.

[0077] The voice acquisition device provided in this application can be a voice acquisition device built into the display device itself, or it can be an external voice acquisition device connected to the display device. For example, it can be the display device's own microphone or an external microphone. This application does not limit the type of voice acquisition device.

[0078] The display device provided in this application can have various implementation forms, such as a smart TV, laser projection device, monitor, electronic bulletin board, electronic table, etc. This application does not limit the type of display device.

[0079] The following is combined Figure 3 The display method provided in the embodiments of this application is described with the controller as the execution subject.

[0080] Figure 3 This is a flowchart illustrating a display method provided in an embodiment of this application, as shown below. Figure 3 As shown, it includes:

[0081] S301. In scenarios involving multiple people in a discussion, the voice recording device performs preliminary recognition of the voices exchanged between users.

[0082] In some embodiments, the controller can process images captured by the camera of the display device, and if the image includes multiple users, it determines that the current scenario is a multi-person discussion.

[0083] In some embodiments, the controller can process the audio collected by the sound acquisition device, and if the audio includes multiple different sound features (such as Mel frequencies), it determines that the current scenario is a discussion involving multiple people.

[0084] In multi-person discussion scenarios, speech recognition technology can be used to initially identify the users' spoken communication and obtain recognition results. For example, the users' spoken communication can be converted into text, and Natural Language Processing (NLP) can be used to recognize the text and obtain recognition results.

[0085] S302. When it is determined that the identification results include tourism-related information, obtain the tourism intentions of each user.

[0086] In some embodiments, if the text recognition results include travel-related information such as "want to go on a beach vacation", "budget of 5,000 yuan", "have time to go out next week", the collected user conversation voice can be accurately recognized to determine each user's travel intention.

[0087] For example, the speech is divided into multiple speech segments; each speech segment includes only one user's speech; the multiple speech segments are clustered to obtain the clustered segments corresponding to each user, and the clustered segments are sorted according to time order to obtain the target speech segments corresponding to each user. Intent recognition is then performed on the target speech segments of each user to obtain the travel intent of each user.

[0088] For example, speech activity detection (VAD) can be used to process the collected speech, removing silent segments and dividing the speech into multiple segments containing speech. In this case, the segmented segments may contain multiple different speakers (users). For each segment, a pre-trained speaker model (such as a deep learning-based model) can be used to process the segment, determining if there is speaker switching within the segment. If so, further segmentation is performed so that each segmented speech segment contains only one user.

[0089] After segmenting into multiple speech segments, audio features can be extracted for each segment. Based on these extracted features, a clustering algorithm is used to cluster the speech segments, resulting in clusters corresponding to each user. Audio features may include Mel-frequency cepstral coefficients, filter bank features, etc.; the clustering algorithm can be k-means, spectral clustering, or deep learning embedding clustering, etc.

[0090] After obtaining the clustered segments corresponding to each user, the clustered segments can be sorted according to their temporal order in the collected speech to obtain the target speech segment corresponding to each user.

[0091] After obtaining the target speech segments, an intent recognition model can be used to identify the intent of each user's target speech segment, thus obtaining each user's travel intent. For example, the target speech segments corresponding to each user can be converted into text, and the text can be processed by a fine-tuned BERT intent classification model to obtain each user's travel intent.

[0092] S303. Generate travel recommendation information based on each user's travel intentions.

[0093] In some embodiments, travel recommendation information may include information such as travel destinations, attraction features, user reviews, itineraries, accommodations, and transportation planning.

[0094] In some embodiments, after determining the travel intentions of each user, algorithms such as ant colony optimization, genetic algorithms, and multi-objective programming can be used to integrate the travel intentions of each user to obtain the optimal travel intention that may satisfy each user, and travel recommendation information can be generated based on the optimal travel intention. For example, artificial intelligence can be used to generate travel recommendation information that satisfies the optimal travel intention.

[0095] S304, Control the display screen to show travel recommendation information.

[0096] In some embodiments, after obtaining travel recommendation information, the travel recommendation information can be converted into a corresponding display video stream and sent to the display screen to show the recommended travel information to the user.

[0097] The display method provided in this application involves: 1) Preliminary recognition of the voice conversations of users collected by a sound acquisition device in a multi-person discussion scenario; 2) Obtaining the travel intentions of each user when the recognition results include tourism-related information; 3) Generating tourism recommendation information based on each user's travel intentions; and 4) Controlling the display screen to show the tourism recommendation information. This method can identify users' travel needs based on their voice conversations and generate travel plans that take into account the personalized needs of family members through recommendation algorithms. Finally, multi-person collaborative decision-making is achieved through large-screen display, significantly improving the efficiency and accuracy of travel planning. This effectively solves the problems of inconvenient information collection, poor interactive experience, and homogenized recommendation results in existing technologies, realizing a new paradigm of intelligent travel planning driven by natural dialogue.

[0098] Below Figure 3 Based on the embodiments shown, the display method provided in this application will be further described.

[0099] Figure 4 A flowchart illustrating a display method provided in an embodiment of this application. Figure 2 ,like Figure 4 As shown, it includes:

[0100] S401. In scenarios involving multiple people in a discussion, the voice recording device performs preliminary recognition of the voices exchanged between users.

[0101] S402. When it is determined that the recognition result includes tourism-related information, the target speech segment corresponding to each user is obtained.

[0102] The specific implementation of steps S401-S402 in this embodiment is similar to that in the above embodiments, and will not be repeated here.

[0103] S403. Perform speech recognition on the target speech segment to obtain the text corresponding to the target speech segment, and perform intent recognition on the text to obtain the first tourism intent.

[0104] In some embodiments, speech-to-text conversion technology can be used to process the target speech segment to obtain the text corresponding to the target speech segment, and an intent recognition model can be used to recognize the intent of the target speech segment corresponding to each user to obtain the explicit first travel intent of each user.

[0105] S404. Based on the tourism knowledge graph, perform intention reasoning on the text to obtain the implicit second tourism intention of the text.

[0106] In some embodiments, to further improve the accuracy of obtaining users' travel intentions, after obtaining the explicit first travel intention, the text corresponding to the target voice segment can be inferred to uncover the user's potential thoughts and obtain the user's implicit second travel intention.

[0107] For example, the text retrieves tourism-related entities; these entities are mapped to corresponding nodes in a tourism knowledge graph; and based on the user's historical tourism information, the user infers a second tourism intent from these nodes.

[0108] Among them, tourism-related entities in the text can refer to places, attractions, activities, etc. mentioned in the text.

[0109] For example, the text is "I want to go to a warm place where I can swim and sunbathe." The corresponding entities are "warm place", "swimming", and "sunbathing".

[0110] After extracting the entities, each entity can be matched and connected with nodes in the tourism knowledge graph, locating a clear and interconnected initial set of nodes in the knowledge graph.

[0111] Starting from a mapped node (such as the "West Lake" node), the user "walks" or expands along the relational edges in the graph, matching and weighting the explored paths and associated nodes with the user's history to output the user's implicit second travel intention.

[0112] The following example illustrates the process of determining a second tourist intention:

[0113] For example, enter the text: "I plan to go to Hangzhou next month. I heard that West Lake and Lingyin Temple are worth seeing."

[0114] Entity recognition: Hangzhou, West Lake, Lingyin Temple, next month;

[0115] Mapping: Map "Hangzhou" to the city node named "Hangzhou City" in the map, and map "West Lake" and "Lingyin Temple" to their respective scenic spot nodes;

[0116] User's historical travel information:

[0117] Past destinations: Suzhou, Nanjing.

[0118] Common activities include visiting historical sites, sampling local snacks, and hiking in nature.

[0119] Consumer preferences: Prefer accommodations with good value for money.

[0120] Reasoning process:

[0121] (1) Path exploration:

[0122] For example: West Lake (belongs to category) → Natural scenery / cultural heritage; Lingyin Temple (belongs to category) → Religious culture / ancient architecture.

[0123] (2) Matching and calculating the relationship between associated nodes and user history:

[0124] Matching Finding 1: The user's history shows a history of "frequently visiting historical sites," and "Lingyin Temple" and "West Lake (cultural landscape)" strongly match this preference. Inference: An implicit intention of the user's trip may be "deep cultural and historical experience," rather than just "sightseeing."

[0125] Matching Finding 2: The user's history shows a preference for "local snacks." In the graph, the "Hangzhou City" node has a "Specialty Food" relationship, linking to nodes such as "Hangzhou Cuisine," "West Lake Vinegar Fish," and "Longjing Shrimp." Inference: The user likely has an implicit intention to "explore local cuisine."

[0126] Related Recommendations: The map shows a "Yanggong Causeway Cycling" activity node near "West Lake," and the user's history includes "Hiking through natural scenery" (both outdoor activities). Inference: The user may also be interested in "Light Outdoor Activities Around the Lake."

[0127] Implied meaning: Recommend cultural and historical sites, Hangzhou cuisine, and outdoor activities around West Lake.

[0128] S405. Merge the first travel intention and the second travel intention to obtain the travel intention of each user.

[0129] In some embodiments, after obtaining the user's explicit first travel intention and implicit second travel intention, they can be merged to obtain the user's true and comprehensive travel needs.

[0130] For example, the first travel intention and the second travel intention are aligned; if the first travel intention and the second travel intention match each other, the second travel intention is supplemented based on the second travel intention to obtain the travel intention; if the first travel intention and the second travel intention conflict, the first travel intention is modified based on the second travel intention to obtain the travel intention.

[0131] For example, the first and second travel intentions can be mapped to the same standard space, and their matching can be compared. If they match, the second travel intention is used to supplement the first travel intention. If there is a conflict, the second travel intention is used to correct the first travel intention.

[0132] The following is a specific example illustrating the fusion process:

[0133] Example 1:

[0134] The primary travel intention is: Hangzhou, weekend, attractions;

[0135] The second travel purpose is: niche travel, in-depth experience, avoiding crowds, exploring shops, and photography;

[0136] Align the two intentions:

[0137] Primary travel intention: {Type: Scenic spot recommendation, Destination: Hangzhou, Core: What to do};

[0138] Second travel intention: {Type: Personalized micro-vacation planning, Destination: Hangzhou, Core: How to travel in a way that suits my unique taste};

[0139] Alignment Analysis:

[0140] The goal is the same: to enjoy a weekend getaway in Hangzhou.

[0141] Relationship assessment: Highly matched and complementary. The second intention does not negate the first intention, but rather adds a personalized filter and a dimension of in-depth planning to the first intention. There are no direct conflicts (such as time or budget conflicts).

[0142] Intent fusion:

[0143] Since the result is a "match", a strategy of "supplementing the first intent based on the second intent" is adopted.

[0144] The ultimate tourism intention is to create personalized, in-depth urban micro-vacation plans, taking the implicit needs of "niche design" and "avoiding crowds" as the core principles for selecting and ranking "attraction recommendations".

[0145] Example 2:

[0146] The primary travel purpose is: business trip, Beijing, Wednesday, to book flight tickets;

[0147] The second travel purpose was: leisure after business trips, weekends, and famous attractions in Beijing;

[0148] Align the two intentions:

[0149] Primary travel intention: {Type: Business travel booking, Destination: Beijing, Core: Booking flights and hotels};

[0150] Second travel intention: {Type: Leisure travel after business trip, Destination: Beijing, Core: May want to do some sightseeing while traveling};

[0151] Alignment Analysis:

[0152] Time correlation: The first intention is to depart next Wednesday, while the second intention is likely to depart after Friday. They are consecutive in time, but their purposes differ.

[0153] Relationship assessment: There is a potential conflict. Business trips typically require hotels close to the meeting venue with convenient transportation, while tourists may prefer accommodations near tourist attractions. Additionally, return flight times may differ (tourism may involve delayed return journeys).

[0154] Intent fusion:

[0155] Since the case is judged to be "conflicting", a strategy of "modifying the first intention based on the second intention" is adopted without violating the first intention.

[0156] The final integrated travel intention: business trip + tourism; departing on Wednesday, meetings on Thursday and Friday mornings, tourism from Friday afternoon to Saturday, and returning on Sunday; booking air tickets (departing on Wednesday and returning on Sunday) and hotels.

[0157] S406. Based on each user's travel intentions, determine each user's travel proposition.

[0158] In some embodiments, during multi-user discussions, different users may have varying degrees of influence on the final travel destination. For example, in family discussions, parents often have a higher influence on the choice of travel destination. Therefore, after obtaining the travel intentions of each user, the travel claims of each user influencing the final travel intention can be calculated.

[0159] For example, the voiceprint features of each user are obtained, and the identity of each user is determined based on the voiceprint features; the weight of the travel intention corresponding to each user is determined according to the identity; and the travel proposition corresponding to each user is determined according to the weight and the travel intention.

[0160] For example, voiceprint features are extracted from target audio segments corresponding to each user, and the extracted voiceprint features are compared with preset voiceprint features to determine the user's identity. For example, the user's identity can be determined as a parent, child, etc.

[0161] After identifying a user, the weights of that user for each dimension of their travel intent can be obtained based on a pre-defined weight mapping relationship. Each dimension of the travel intent can refer to different options within the travel intent, such as destination, type of travel, and travel budget.

[0162] After obtaining the weights of each dimension, the dimensions of travel intent can be labeled based on the weights to obtain the travel proposition of each user. For example, user A's travel proposition is: {Type = Attraction recommendation, weight = 0.8; Destination = Hangzhou, weight = 0.5; Budget = 5000 yuan, weight = 0.3}.

[0163] S407. Construct a tourism recommendation network based on the tourism propositions of each user, and determine the group tourism consensus vector of each user based on the tourism recommendation network.

[0164] In some embodiments, after obtaining the travel claims of each user, the rebuttal or support relationships between the claims can be extracted, and a travel recommendation network can be constructed based on the claims and the rebuttal or support relationships. In this travel recommendation network, nodes represent claims, edges represent rebuttal or support relationships, and each claim node has the following attributes: dimension, weight, and user identifier.

[0165] After constructing the travel recommendation network, a graph neural network (GNN) can be used to process it, integrating all claims and relationships, updating the representation of each claim, and then calculating the group travel consensus vector for each user. The group travel consensus vector can include consensus representations for each dimension (such as the value or distribution of each dimension) and the degree of controversy for each dimension (a high degree of controversy indicates significant disagreement on claims within that dimension).

[0166] For example, suppose there are n claims (i.e., n users), construct a graph G=(V,E). Each node v_i represents a claim, and node features include: dimension (e.g., destination), value (e.g., beach), weight (0.8), user ID, etc. Edge e_ij represents the relationship between claim i and claim j, and the relationship type can be supportive or refutative.

[0167] The graph is processed using Graph Convolutional Networks (GCN), Graph Attention Networks (GAT), or Relational Graph Convolutional Networks (R-GCN). Assuming each node v_i has an initial feature vector h_i^0, after propagation through multiple layers of GNNs, the final node representation h_i^L is obtained. All claim node representations (h_i^L) belonging to this dimension are collected and aggregated using a weighted average (weights being node weights) or an attention mechanism to calculate the consensus representation for this dimension. The degree of controversy in this dimension is calculated by evaluating the variance between node representations or by calculating the attention weight entropy during weighted aggregation.

[0168] For example, suppose dimension d has m claim nodes, represented as h_1, h_2, ..., h_m, with corresponding weights w_1, w_2, ..., w_m.

[0169] The consensus statement is: c_d = sum(w_i * h_i) / sum(w_i)

[0170] Degree of controversy: This can be calculated by weighing the variance or by using entropy. For example, first normalize the weights, then calculate the entropy: H_d = -sum(p_i * log(p_i)), where p_i = w_i / sum(w_i). The higher the entropy, the higher the degree of controversy.

[0171] S408. Generate tourism recommendation information based on the group tourism consensus vector.

[0172] In some embodiments, after obtaining the group tourism consensus vector, tourism recommendation information can be generated based on the consensus representation and degree of controversy in the group tourism consensus vector.

[0173] For example, the degree of controversy in each dimension of the group tourism consensus vector is determined; when the degree of controversy is less than or equal to a preset value, a first recommended option is generated based on the corresponding dimension consensus; when the degree of controversy is greater than the preset value, multiple second recommended options are generated based on the corresponding dimension controversy; the first recommended option and the multiple second recommended options are merged to obtain tourism recommendation information.

[0174] For example, for low-controversy dimensions (the degree of controversy is lower than the preset value), the value corresponding to the consensus representation is directly used as the recommendation input to obtain the first recommended option (for example, for the destination dimension, if the consensus representation vector is closest to the embedding vector of "seaside", then seaside is selected).

[0175] For highly controversial dimensions (those with a level of controversy higher than the preset value), multiple candidate values ​​are retained (e.g., for the destination dimension, there are both seaside and mountain areas). When generating recommendations, options that take into account these candidate values ​​are provided (e.g., recommending a scenic spot with both sea and mountain views, or recommending itineraries to two different destinations).

[0176] S409. At the target time, control the display screen to show travel recommendation information.

[0177] In some embodiments, when displaying travel recommendations to users, in order to avoid the travel recommendations interfering with the users, the controller can display the travel recommendations at an appropriate time, such as when users are watching TV and discussing travel destinations, to avoid suddenly displaying travel recommendations and interrupting the users' viewing process.

[0178] For example, the current environmental state and display content are obtained; based on the environmental state and display content, the target time for displaying travel recommendation information is determined; at the target time, the display screen is controlled to display the travel recommendation information.

[0179] The environmental state can include whether the environment is quiet or active (conversation, laughter), the current time, etc. The displayed content can include currently displayed TV content, advertisements, end credits, etc.

[0180] The system can process the environmental state and displayed content based on a predefined set of rules or a pre-trained machine learning model to determine the target time for displaying travel recommendations. For example, travel recommendations could be displayed during the end credits and when users are interacting.

[0181] In some embodiments, after displaying travel recommendations, user feedback instructions can be obtained; and the travel recommendations can be adjusted based on the feedback instructions.

[0182] For example, users can adjust their travel budget via voice input. The controller receives the adjusted budget, updates the travel recommendations based on the new budget, and controls the display screen to show the updated travel recommendations.

[0183] In summary, the display method provided in this application automates the entire process from natural family conversations to travel plan generation. Multi-channel voice acquisition and sound source separation technologies ensure accurate capture of each family member's voice needs even in complex acoustic environments. Semantic analysis technology identifies implicit needs in the conversation, and knowledge graph technology integrates scattered information. A recommendation algorithm based on multi-dimensional feature fusion generates travel plans that take into account the personalized needs of family members. Finally, a large-screen display enables collaborative decision-making among multiple people, significantly improving the efficiency and accuracy of travel planning.

[0184] Based on the above embodiments, this application also provides a display device.

[0185] Figure 5 This is a schematic diagram of the structure of the display device 50 provided in the embodiments of this application, as shown below. Figure 5 As shown, it includes:

[0186] The acquisition module 501 is used to acquire the voice of users communicating in a multi-person discussion scenario and to perform content recognition on the voice.

[0187] The acquisition module 502 is used to acquire the travel intent of each user when the content recognition result is determined to include travel-related information.

[0188] The processing module 503 is used to generate travel recommendation information based on the travel intentions of each user.

[0189] The control module 504 is used to control the display screen to show travel recommendation information.

[0190] In some embodiments, the acquisition module 502 is used to divide the speech into multiple speech segments; each speech segment includes only the speech of one user; the multiple speech segments are clustered to obtain the clustered segments corresponding to each user, and the clustered segments are sorted according to time order to obtain the target speech segments corresponding to each user. Intent recognition is performed on the target speech segments of each user to obtain the travel intent of each user.

[0191] In some embodiments, the acquisition module 502 is used to perform speech recognition on the target speech segment to obtain the text corresponding to the target speech segment; perform intent recognition on the text to obtain a first tourism intent; perform intent reasoning on the text based on a tourism knowledge graph to obtain a second tourism intent implied in the text; and fuse the first tourism intent and the second tourism intent to obtain a tourism intent.

[0192] In some embodiments, the acquisition module 502 is used to acquire entities related to tourism in the text; map the entities to the corresponding nodes in the tourism knowledge graph; and, starting from the nodes, reason based on the user's historical tourism information to obtain a second tourism intention.

[0193] In some embodiments, the acquisition module 502 is used to align the first travel intention and the second travel intention; if the first travel intention and the second travel intention match each other, the second travel intention is supplemented based on the second travel intention to obtain the travel intention; if the first travel intention and the second travel intention conflict, the first travel intention is corrected based on the second travel intention to obtain the travel intention.

[0194] In some embodiments, the processing module 503 is configured to acquire the voiceprint features of each user, determine the identity of each user based on the voiceprint features, determine the weight of the travel intention corresponding to each user based on the identity, determine the travel proposition corresponding to each user based on the weight, travel intention and identity, construct a travel recommendation network based on the travel proposition corresponding to each user, and determine the group travel consensus vector of each user based on the travel recommendation network, and generate travel recommendation information based on the group travel consensus vector.

[0195] In some embodiments, the processing module 503 is used to determine the degree of controversy in each dimension of the group tourism consensus vector; when the degree of controversy is less than or equal to a preset value, a first recommended option is generated based on the corresponding dimension consensus; when the degree of controversy is greater than the preset value, multiple second recommended options are generated based on the corresponding dimension controversy; and the first recommended option and multiple second recommended options are merged to obtain tourism recommendation information.

[0196] In some embodiments, the control module 504 is configured to acquire the current environmental state and display content; determine the target time for displaying travel recommendation information based on the environmental state and display content; and control the display screen to display the travel recommendation information at the target time.

[0197] In some embodiments, the processing module 503 is used to obtain user feedback instructions on travel recommendation information and adjust the travel recommendation information based on the feedback instructions.

[0198] The display device provided in this application embodiment can perform the display function shown in any of the above embodiments, and its principle and technical effect are similar, so it will not be described again here.

[0199] This application also provides an electronic device.

[0200] Figure 6 This is a schematic diagram of the structure of the electronic device 60 provided in the embodiments of this application, such as... Figure 6 As shown, the electronic device may include: a transceiver 601, a processor 602, and a memory 603. The electronic device may be a controller as described in any of the above embodiments.

[0201] The processor 602 executes computer execution instructions stored in the memory, causing the processor 602 to perform the scheme in the above embodiments. The processor 602 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0202] The memory 603 is connected to the processor 602 via the system bus and completes communication between them. The memory 603 is used to store computer program instructions.

[0203] Transceiver 601 can perform the functions of receiving and sending data and instructions.

[0204] Optionally, the electronic device 60 may also include a communication interface 604, which allows communication and interaction with external or internal devices via the communication interface 603. External devices may be, for example, client devices (e.g., mobile phones, tablets). In specific implementations, if the communication interface 604, memory 603, and processor 602 are implemented independently, they can be interconnected via a bus to complete communication with each other.

[0205] The system bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. Transceivers are used to enable communication between database access devices and other computers (e.g., clients, read-write libraries, and read-only libraries). Memory may include random access memory (RAM) and may also include non-volatile memory.

[0206] Optionally, in a specific implementation, if the communication interface 604, memory 603, and processor 602 are integrated on a single chip, then the communication interface 604, memory 603, and processor 602 can communicate through an internal interface.

[0207] This application also provides a chip for executing instructions, which is used to execute the technical solutions of the methods described in the above embodiments.

[0208] This application also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the technical solution of the above method embodiment. Its implementation principle and technical effect are similar, and will not be repeated here.

[0209] In one possible implementation, a computer-readable medium may include random access memory (RAM), read-only memory (ROM), compact discread-only memory (CD-ROM) or other optical disc storage, disk storage or other magnetic storage devices, or any other medium targeted to carry or to store the required program code in the form of instructions or data structures, and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. As used herein, disks and optical discs include optical discs, laser discs, optical discs, Digital Versatile Discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs optically reproduce data using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0210] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the technical solution of the above method embodiments. Its implementation principle and technical effects are similar, and will not be repeated here.

[0211] In the specific implementation of the aforementioned terminal device or server, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly manifested as execution by a hardware processor, or execution by a combination of hardware and software modules within the processor.

[0212] Those skilled in the art will understand that all or part of the steps in any of the above method embodiments can be implemented by hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium, and when the program is executed, all or part of the steps in the above method embodiments are performed.

[0213] If the technical solution of this application is implemented in software form and sold or used as a product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the technical solution of this application can be embodied in the form of a software product, which is stored in a storage medium and includes a computer program or several instructions. This computer software product enables a computer device (which may be a personal computer, server, network device, or similar electronic device) to execute all or part of the steps of the methods in the embodiments of this application.

[0214] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0215] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0216] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.

[0217] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.

[0218] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.

[0219] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0220] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.

[0221] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A display device, characterized in that, The display device includes: The display screen is configured to display images. A sound acquisition device, configured to acquire sound; A controller connected to the display screen is configured to: In scenarios involving multiple people discussing, the voice recording device is used to perform preliminary recognition of the users' conversations. When the identification results include tourism-related information, obtain the tourism intent of each user; Based on the travel intentions of each user, travel recommendation information is generated; Control the display screen to show the travel recommendation information.

2. The display device according to claim 1, characterized in that, The controller is configured to: The speech is divided into multiple speech segments; each speech segment includes only the speech of one user. The multiple speech segments are clustered to obtain the clustered segments corresponding to each user, and the clustered segments are sorted according to the time order to obtain the target speech segments corresponding to each user. The travel intentions of each user are obtained by performing intent recognition on the target speech segments of each user.

3. The display device according to claim 2, characterized in that, The controller is configured to: The target speech segment is subjected to speech recognition to obtain the text corresponding to the target speech segment; The text is subjected to intent recognition to obtain the first travel intent; Intent reasoning is performed on the text based on a tourism knowledge graph to obtain the implicit second tourism intent of the text. The first travel intention and the second travel intention are merged to obtain the travel intention.

4. The display device according to claim 3, characterized in that, The controller is configured to: Obtain the entities related to tourism from the text; The entity is mapped to the corresponding node in the tourism knowledge graph, and the second tourism intention is obtained by reasoning based on the user's historical tourism information starting from the node.

5. The display device according to claim 4, characterized in that, The controller is configured to: Align the first travel intention and the second travel intention; If the first travel intention and the second travel intention match each other, then the second travel intention is supplemented based on the second travel intention to obtain the travel intention; If there is a conflict between the first travel intention and the second travel intention, the first travel intention is modified based on the second travel intention to obtain the travel intention.

6. The display device according to claim 5, characterized in that, The controller is configured to: Obtain the voiceprint features of each user, and determine the identity of each user based on the voiceprint features; Based on the identity identifier, determine the weight of each user's corresponding travel intention; Based on the weights and the travel intentions, the travel propositions corresponding to each user are determined. A tourism recommendation network is constructed based on the tourism propositions of each user, and a group tourism consensus vector for each user is determined based on the tourism recommendation network. The tourism recommendation information is generated based on the group tourism consensus vector.

7. The display device according to claim 6, characterized in that, The controller is configured to: Determine the degree of controversy for each dimension in the aforementioned group tourism consensus vector; When the degree of controversy is less than or equal to a preset value, a first recommended option is generated based on the corresponding dimensional consensus. When the degree of controversy exceeds the preset value, multiple second recommendation options are generated based on the corresponding dimension of controversy. The first recommended option and multiple second recommended options are combined to obtain the travel recommendation information.

8. The display device according to any one of claims 1-7, characterized in that, The controller is configured to: Get the current environment status and displayed content; Based on the environmental conditions and the displayed content, determine the target time for displaying the tourism recommendation information; At the target time, the display screen is controlled to show the travel recommendation information.

9. The display device according to claim 8, characterized in that, The controller is configured to: Obtain user feedback instructions regarding the travel recommendation information; The travel recommendation information is adjusted based on the feedback instructions.

10. A display method, characterized in that, include: In scenarios involving multiple people in a discussion, the system collects the voice recordings of users' conversations and performs content recognition on the voice recordings. When the content recognition results are determined to include tourism-related information, the tourism intent of each user is obtained. Based on the travel intentions of each user, travel recommendation information is generated; Control the display screen to show the travel recommendation information.

11. A computer-readable storage medium, characterized in that, It stores a computer program, which is executed by a processor to implement the method of claim 10.

12. A computer program product, characterized in that, It includes a computer program that, when executed by the controller, implements the method of claim 10.