Comment generation method, comment display method and related device
By obtaining the attribute information of images uploaded by users and target points of interest, and combining the multimodal large model and RAG knowledge base to generate high-quality review prompt information, the problems of insufficient depth and diversity in existing review generation systems are solved, and more efficient and personalized review generation is achieved.
Patent Information
- Application Number
- CN202510954722.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-16
AI Technical Summary
Existing review generation systems rely on text information input by users to generate content, which results in insufficient depth and diversity in the generated reviews, resulting in limited value of the review content and affecting user experience.
By obtaining images uploaded by users, combining the attribute information of target points of interest and the user's emotional tendencies, and using a pre-trained comment generation model to generate candidate comments, including the use of a multimodal large model and RAG knowledge base, high-quality comment prompt information is generated.
It reduces the difficulty for users to create comments, improves the efficiency and quality of comment generation, provides users with a more convenient interactive experience, and generates comments that are more realistic and personalized.
Smart Images

Figure CN120654676A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information service technology, and in particular to a comment generation method, a comment display method and related devices. Background Art
[0002] Reviews, as a key means of online word-of-mouth communication, are crucial for users to understand the pros and cons of products and services and make informed choices. However, the review creation process often requires users to organize their words, format and edit them. This process is not only complex and time-consuming, but can easily lead some users to abandon their review due to inconvenience.
[0003] With the continuous development of artificial intelligence technology, a number of review generation systems have emerged, aiming to help users automatically generate review content. However, existing review generation systems mostly rely solely on user-entered text information to generate content, which has many limitations. These systems often simply expand or style the input text during the generation process, rarely adding additional information. This results in insufficient depth and diversity in the generated reviews, and their value is limited. In some cases, the generated reviews may even be redundant and contain excessive clichés, affecting the user experience. Summary of the Invention
[0004] This application provides a comment generation method, a comment display method and related devices, aiming to generate high-quality comments for users and improve user experience.
[0005] This application provides the following solutions:
[0006] According to a first aspect, a comment generation method is provided, the method comprising:
[0007] Obtaining a comment generation request for a target point of interest, the comment generation request including at least one image uploaded by a user;
[0008] Determining attribute information of the target point of interest and the target emotional tendency of the user toward the target point of interest;
[0009] Generate comment prompt information based on the attribute information of the target interest point and the target emotional tendency;
[0010] The comment prompt information and the image are input into a pre-trained comment generation model to generate candidate comments that meet the emotional tendency of the comments.
[0011] According to a second aspect, a comment display method is provided, the method comprising:
[0012] Displaying a comment posting page for the target point of interest, wherein the comment posting page includes an image upload area and a comment generation control;
[0013] In response to the image upload area detecting that the user has uploaded at least one image and the comment generation control being triggered, the generated candidate comments are displayed, the candidate comments are generated based on the image and comment prompt information, and the comment prompt information is generated based on the attribute information of the target point of interest and the user's target emotional tendency towards the target point of interest.
[0014] According to a third aspect, a comment generation device is provided, the device comprising:
[0015] a request acquiring unit configured to acquire a comment generation request for a target point of interest, wherein the comment generation request includes at least one image uploaded by a user;
[0016] an information determining unit, configured to determine attribute information of the target point of interest and a target emotional tendency of the user toward the target point of interest;
[0017] a prompt generating unit configured to generate comment prompt information based on the attribute information of the target interest point and the target emotional tendency;
[0018] The comment generation unit is configured to input the comment prompt information and the image into a pre-trained comment generation model to generate candidate comments that meet the comment sentiment tendency.
[0019] According to a fourth aspect, a comment display device is provided, comprising:
[0020] A page display unit is configured to display a comment publishing page of the target point of interest, wherein the comment publishing page includes an image upload area and a comment generation control;
[0021] A comment generation unit is configured to display generated candidate comments in response to the image upload area monitoring that the user has uploaded at least one image and the comment generation control is triggered, wherein the candidate comments are generated based on the image and comment prompt information, and the comment prompt information is generated based on the attribute information of the target point of interest and the user's target emotional tendency towards the target point of interest.
[0022] According to a fifth aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the method described in any one of the first aspect or the second aspect.
[0023] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0024] This embodiment of the present application obtains a comment generation request for a target point of interest (POI), which includes a user-uploaded image. It then generates comment prompt information based on the determined attribute information of the target POI and the user's target emotional inclination toward the target POI. The comment prompt information and the user-uploaded image are then fed into a pre-trained comment generation model to generate candidate comments for the target POI. Compared to existing technologies, this method only requires the user to upload an image to generate corresponding candidate comments, effectively reducing the difficulty of user comment creation, improving the efficiency and quality of comment generation, and providing users with a more convenient interactive experience.
[0025] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0027] Figure 1 is a system architecture diagram applicable to the embodiments of the present application;
[0028] Figure 2 Flowchart of the comment generation method provided by the embodiment of the present application;
[0029] Figure 3 A flowchart of a method for displaying comments provided in an embodiment of the present application;
[0030] Figure 4 A schematic diagram of a comment generation page provided in an embodiment of the present application;
[0031] Figure 5 A flowchart of a method for training a review generation model provided in an embodiment of the present application;
[0032] Figure 6 A flowchart of the method for generating comments provided in an embodiment of the present application;
[0033] Figure 7 A schematic block diagram of a review generating apparatus provided in an embodiment of the present application;
[0034] Figure 8 A schematic block diagram of a comment display device provided in an embodiment of the present application;
[0035] Figure 9 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0037] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "an", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0038] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0039] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0040] Existing review generation systems mostly rely solely on user-entered textual information for content generation, which presents numerous limitations. These systems often simply expand or style the input text, rarely adding additional information. This results in insufficient depth and diversity in the generated reviews, limiting their value. In some cases, the generated reviews may even be overly redundant and clichéd, impacting the user experience.
[0041] In view of this, the present application provides a new approach. To facilitate understanding of the present application, the system architecture on which the present application is based is first described. Figure 1 An exemplary system architecture to which the embodiments of the present application can be applied is shown. Figure 1 As shown in , the system architecture may include: a user end, a terminal device and a server end.
[0042] Among them, the user end is set on the terminal device. The user end involved in the embodiment of the present application can be a client running on the terminal device, a small program or a Web application running through a browser, etc.
[0043] Terminal devices may include, but are not limited to, smart mobile terminals, wearable devices, PCs (Personal Computers), smart home devices, and the like. Smart mobile devices may include, for example, mobile phones, tablets, laptops, PDAs (Personal Digital Assistants), and internet-connected car terminals. Wearable devices may include, for example, smart watches, smart glasses, smart bracelets, VR (Virtual Reality) devices, AR (Augmented Reality) devices, mixed reality devices (i.e., devices that support both virtual reality and augmented reality), and the like. Smart home devices may include, for example, smart TVs and smart refrigerators with displays.
[0044] The user side can interact with the server side through the network, post comments, or view comments posted by others.
[0045] The server side can be the backend server for the client, providing backend services for the client, such as generating candidate reviews for the user. The server side can be a single server, a server cluster consisting of multiple servers, or a cloud server. A cloud server, also known as a cloud computing server or cloud host, is a hosting product within the cloud computing service ecosystem. It addresses the management difficulties and limited scalability of traditional physical hosting and virtual private server (VPS) services.
[0046] It should be understood that Figure 1 The number of user terminals, terminal devices and server terminals in the embodiment is merely illustrative. Any number of user terminals, terminal devices and server terminals may be provided according to implementation requirements.
[0047] Figure 2 This is a flow chart of the comment generation method provided in the embodiment of the present application. The method can be Figure 1 The server side execution in the system shown. Figure 2 As shown in , the method may include the following steps:
[0048] Step 201: Obtain a comment generation request for a target point of interest, where the comment generation request includes at least one image uploaded by a user;
[0049] Step 202: Determine the attribute information of the target point of interest and the user's target emotional tendency toward the target point of interest;
[0050] Step 203: Generate comment prompt information based on the attribute information and target emotional tendency of the target interest point;
[0051] Step 204: Input the comment prompt information and the image into a pre-trained comment generation model to generate candidate comments corresponding to the target interest point.
[0052] As can be seen, the present embodiment obtains a comment generation request for a target POI, which includes a user-uploaded image. It then generates comment prompt information based on the determined attribute information of the target POI and the user's target emotional inclination toward the target POI. The comment prompt information and the user-uploaded image are then fed into a pre-trained comment generation model to generate candidate comments corresponding to the target POI. Compared to existing technologies, this method only requires the user to upload an image to generate the corresponding candidate comments, effectively reducing the difficulty of user comment creation, improving the efficiency and quality of comment generation, and providing users with a more convenient interactive experience.
[0053] The steps in the above process and the effects that can be further produced are described in detail below with reference to embodiments.
[0054] The above step 201, namely, "obtaining a comment generation request for a target point of interest, where the comment generation request includes at least one image uploaded by a user," is described in detail below with reference to an embodiment.
[0055] A target point of interest (POI) is a specific geographical entity or place that a user wishes to review. These POIs can be restaurants, hotels, scenic spots, parks, shops, and any other places that are attractive to the user.
[0056] Before submitting a review generation request, users can upload at least one image to the review publishing page for the target point of interest. These images provide rich visual information for review generation, helping to generate more realistic and attractive reviews. User-uploaded images may include: environmental images, i.e., photos of the interior or exterior environment of the target point of interest, such as the dining environment of a restaurant, the lobby or room of a hotel, the scenery of a tourist attraction, etc.; product or service images, such as restaurant dishes, hotel facilities, and the characteristic landscape of a tourist attraction, etc.; user experience images, i.e., photos of the user's personal experience at the target point of interest, such as photos with friends in a restaurant, souvenirs at a tourist attraction, etc. These images not only provide an intuitive visual reference for review generation, but also help the review generation model better understand the user's experience and feelings.
[0057] It should be noted that when a user uploads images, the number of images can be one or more, and this embodiment of the application does not impose any restrictions on this.
[0058] In addition to the user-uploaded image, the comment generation request may also include the target point of interest (POI) identification information and the user's identity information. The target point of interest (POI) identification information is used to identify and locate the POI, and typically includes, but is not limited to, the POI ID. The user's identity information is used to identify and distinguish different users, and typically includes, but is not limited to, the user ID.
[0059] Furthermore, when a user submits a comment generation request, the user may also include comment information on the target point of interest, which may include text comment information and / or rating information, etc. Such information may also provide a basis for generating candidate comments later.
[0060] In actual applications, when a user wants to comment on a target point of interest, taking a map application as an example, the user can click on the target point of interest in the map application to enter the comment publishing page of the target point of interest, and submit a comment generation request by clicking the comment generation control on the comment publishing page. The specific implementation process will be described in detail later.
[0061] The above step 202, namely "determining the attribute information of the target point of interest and the target emotional tendency of the user towards the target point of interest", is described in detail below in conjunction with an embodiment.
[0062] In this embodiment of the present application, after receiving a request to generate a review for a target POI, the server first determines the attribute information of the target POI and the user's target sentiment toward the target POI. The attribute information of the target POI may include, but is not limited to, the name, category, address, business hours, facility type, business district location, recommended items, and the like.
[0063] The attribute information of the target point of interest can be determined based on the identification information of the target point of interest (such as the POI ID of the target point of interest) included in the comment generation request. Specifically, the attribute information of the target point of interest can be obtained in the following ways:
[0064] (1) Obtaining basic attribute information of points of interest from a pre-stored database
[0065] The basic attribute information of the target point of interest corresponding to the POIID is obtained from a database that pre-stores the basic attribute information of the points of interest through the POIID, and the attribute information of the target point of interest is determined based on these basic attribute information. Among them, the basic attribute information of the target point of interest is a direct description of the basic characteristics of the target point of interest, such as name, category, geographical location, etc., and the attribute information, in addition to including the basic attribute information, can also be based on the basic attribute information, expanded with scenario-based and detailed descriptive information. For example, the geographical location in the basic attribute information is Pudong New Area, Shanghai. Based on such basic attribute information, the corresponding attribute information is further expanded to "2.7 kilometers from the city center" and "nearby Lujiazui business district".
[0066] (2) Real-time acquisition through third-party data interface
[0067] By accessing third-party data interfaces on external platforms, dynamic attribute information of target points of interest can be obtained in real time, such as temporary event information at scenic spots and new product recommendations at restaurants. This approach ensures the timeliness and richness of attribute information.
[0068] (3) Extracting historical review information from the target point of interest
[0069] The text model is used to analyze and extract historical reviews of the target POI, extracting valid attribute information from it. For example, from the historical review information "the dishes are generous in portion", the dish attributes of the restaurant (i.e., the target POI) are extracted as one of the restaurant's attribute information.
[0070] (4) Retrieval from a pre-built Retrieval-Augmented Generation (RAG) knowledge base
[0071] Quickly retrieve and obtain relevant attribute information of target points of interest through the pre-built RAG knowledge base. This knowledge base integrates multi-source data to provide richer and more accurate attribute descriptions, further improving the accuracy and personalization of review generation.
[0072] Through the above methods, the attribute information of the target point of interest can be obtained comprehensively and accurately.
[0073] The target emotional tendency of the user towards the target point of interest refers to the attitude tendency shown by the user when commenting on the target point of interest, which can usually be divided into positive (i.e., positive attitude), neutral (i.e., neutral attitude) and negative (i.e., negative attitude) types. In an embodiment of the present application, if the user does not actively input comment information, a pre-set default emotional tendency can be used as the user's target emotional tendency towards the target point of interest. For example, the emotional tendency can be represented by a quantitative indicator, and a scoring interval of 1 to 5 points can be set, where the lower the score, the more negative the attitude, and the higher the score, the more positive the attitude. In an embodiment of the present application, the default emotional tendency can be set to 4 or 5 points to reflect a more positive attitude tendency.
[0074] When the user actively inputs comment information about the target point of interest, the target sentiment tendency can be determined based on the comment information, which may include the following situations:
[0075] When a user enters a text review, natural language processing techniques (such as sentiment analysis models or text classification algorithms) can be used to perform semantic analysis on the review, automatically identifying and extracting the user's expressed sentiment, thereby determining the target sentiment. For example, if the user's review reads "the environment is great, the service is attentive," the target sentiment can be determined to be positive; if the review reads "the food is average, the price is high," the target sentiment can be determined to be neutral or negative.
[0076] When a user enters a rating for a target POI, the rating can be directly mapped to the corresponding target sentiment. For example, if the rating range is 1 to 5, a rating of 4 or 5 can be mapped to a positive sentiment, a rating of 3 to a neutral sentiment, and a rating of 1 or 2 to a negative sentiment.
[0077] The above method can flexibly and accurately determine the user's target emotional tendency towards the target interest point, thereby generating comment content that is more in line with the user's true intention and improving user experience and interaction effects.
[0078] It should be noted that, in the embodiment of the present application, there is no specific restriction on the execution order of determining the attribute information of the target interest point and determining the target emotional tendency of the user towards the target interest point in step 202. In practical applications, it can be flexibly set according to specific needs.
[0079] The above step 203, namely "generating comment prompt information based on the attribute information and target emotional tendency of the target interest point", is described in detail below in conjunction with an embodiment.
[0080] After determining the attribute information of the target POI and the user's target emotional inclination towards the target POI, comment prompt information can be generated based on this information. Specifically, target knowledge items related to the attribute information of the target POI and that match the target emotional inclination can be filtered out from a pre-built RAG knowledge base. Then, based on the attribute information of the target POI and the filtered target knowledge items, comment prompt information can be generated according to a preset structured template.
[0081] In the embodiment of the present application, the RAG knowledge base is a pre-built knowledge base, and its construction process is as follows:
[0082] The pre-trained text model is used to analyze the historical comment information of the points of interest and extract knowledge items that meet the preset conditions. The knowledge items include the content of the knowledge points and the corresponding emotional tendencies; the extracted knowledge items are then stored in the RAG knowledge base.
[0083] The content of the knowledge points may include but is not limited to the features, advantages, disadvantages, user evaluations, etc. of the points of interest, and the corresponding emotional tendencies may include but are not limited to positive, neutral, negative, etc.
[0084] To ensure that historical reviews of POIs contain knowledge items that can provide users with decision-making information, a pre-built template for extracting knowledge items can be used as a precondition to guide the large text model's screening. This historical review information can be obtained from multiple channels, such as online review platforms, social media, and user feedback.
[0085] Optionally, the constructed prompt word template for knowledge entry extraction may include the following content:
[0086] Content requirements: Knowledge points must contain valid information, be independent and complete, and be able to provide assistance to other users. Knowledge points must be no less than 5 words long and avoid being too brief.
[0087] Emotional tendency: The emotional tendency of knowledge point content is divided into three types: "positive", "neutral" and "negative".
[0088] Expression requirements: The expression of knowledge points needs to be concise. Appropriate restatement can be used in the language, but it should not contain content that is overly emotional, advertising, or marketing, such as "highly recommended" or "very good."
[0089] Relevance: The knowledge point content cannot contain descriptions that are completely unrelated to the main location information.
[0090] As an example, the historical review information for a point of interest is as follows:
[0091] Location: Jiulongpo District, Chongqing
[0092] Scenery: Going in June will bring you lush vegetation and great views. It's recommended to go in the morning or evening, as the furry animals are taking naps in the afternoon. Recommended activities: Red Panda Pavilion, White Tiger, and the monkey named Ping An. The name of the most approachable animal: Costa's straight-line carp.
[0093] Zoological classification: Actinopterygii, Cyprinidae, Cyprinidae, Cyprinidae
[0094] Diet: Feeds on algae, aquatic plants, aquatic insects, small crustaceans, etc.
[0095] Based on the above-constructed prompt word template for knowledge item extraction, the following knowledge items can be extracted from the above-mentioned historical review information:
[0096] Knowledge item 1: Location is in Jiulongpo District, Chongqing (knowledge point content); neutral (emotional tendency)
[0097] Knowledge item 2: Scenery: The vegetation is lush in June, and the scenery is just right (knowledge point content); positive (emotional tendency)
[0098] Knowledge item 3: It is recommended to go in the morning or evening (knowledge point content); neutral (emotional tendency)
[0099] Knowledge item 4: The furries are taking a nap in the afternoon (knowledge point content); neutral (emotional tendency)
[0100] Knowledge item 5: Project recommendation: Red Panda House and White Tiger House (knowledge point content); neutral (emotional tendency)
[0101] Knowledge item 6: The little monkey named Ping An is very friendly (knowledge point content); positive (emotional tendency)
[0102] The RAG knowledge base constructed through the above process contains knowledge items that not only include specific knowledge point content, such as users' specific evaluations and feelings about points of interest, but also clearly mark the emotional tendencies corresponding to each knowledge point content. Therefore, when filtering based on the attribute information of the target point of interest and the user's target emotional tendencies towards the point of interest, it is possible to achieve accurate matching in the RAG knowledge base and effectively extract target knowledge items that are relevant in content and consistent in emotion. This mechanism ensures that the information cited in the review generation process is both consistent with the actual situation of the target point of interest and meets the user's emotional expression needs, thereby significantly improving the relevance and personalization of the review content.
[0103] The construction and use of the RAG knowledge base can effectively reduce the hallucination phenomenon of the review generation model and improve the quality and credibility of the generated reviews.
[0104] Furthermore, when generating comment prompt information, in addition to the attribute information of the target POI and the user's target emotional tendency towards the target POI, the user's profile information can also be used. The user's profile information refers to comprehensive data used to describe the user's characteristics, preferences, and behavior patterns, and may include but is not limited to the user's preference information (such as the types of POIs frequently visited, consumption habits, language style preferences, etc.), the user's evaluation information (such as historical rating distribution, high-frequency words in review topics, etc.), and the user's emotional state information (such as the emotional polarity of historical reviews, the emotional tendency corresponding to the ratings, etc.).
[0105] User profile information can be obtained in the following ways:
[0106] Obtain the user profile data from a database that has pre-stored user profile data. For example, retrieve the pre-stored profile data by user ID.
[0107] Obtained from the user's historical comments. Natural language processing technology can be used to analyze historical comment information and extract profile information such as high-frequency words in the comment topic and sentiment tendencies.
[0108] Obtained from third-party data interfaces. Obtain the user's public behavior data by accessing external platforms.
[0109] Since user portrait information can accurately reflect their personalized needs and characteristics, the comment prompt information generated based on this information is more targeted and personalized, thereby improving the satisfaction and practicality of the generated comments.
[0110] Optionally, comment prompts can be generated based on the target POI's attribute information, the user's profile information, and the user's emotional inclination toward the target POI. This comprehensive integration ensures that the generated comment prompts not only include the basic characteristics of the target POI but also reflect the user's personalized preferences and subjective feelings. Through this multi-dimensional information fusion, the generated comment prompts will be more closely aligned with the user's actual needs and emotional expressions, thereby guiding the generation of more personalized, authentic, and engaging comment content, significantly improving the helpfulness and practicality of comments for users.
[0111] In order to enable the review generation model to fully understand the intention of review generation, the embodiment of the present application structures the information elements of review generation. By integrating these elements into a preset structured template, high-quality, personalized review prompt information can be generated.
[0112] The structured template includes attribute information of the point of interest and knowledge base information. In addition, it may also include but is not limited to at least one element of language style, topic type and comment template.
[0113] The language style is flexible. The server can pre-set multiple language styles, such as concise, approachable, lively, and authoritative. When generating comment prompts, it can randomly select one of these pre-set language styles to ensure the diversity and interest of the comments. Alternatively, it can select a corresponding language style from the pre-set language styles based on the user's profile information. This allows for precise matching of the user's expression habits and emotional tendencies, thereby enhancing the personalization and relevance of the generated comments.
[0114] Topic types refer to thematic dimensions commonly found in review descriptions. These dimensions are determined based on the type of point of interest and the user's focus. For example, for tourist attractions, topic types might include scenery, history, culture, facilities, transportation, accommodation, room rates, business hours, and ticket prices. By specifying topic types, generated reviews can be more specific and targeted.
[0115] A review template is a structured expression of a review, guiding the review generation model to generate template structures with different characteristics. It also offers flexibility. The server can pre-set a variety of review templates, such as Xiaohongshu-style templates and Dianping-style templates. These typically include categories and some semantic emoji symbols to enhance the appeal and readability of the review. When generating review prompts, one of the pre-set review templates can be randomly selected, or a corresponding review template can be selected from the pre-set review templates based on the attribute information of the target point of interest.
[0116] The knowledge base information is determined based on the target knowledge items retrieved from the RAG knowledge base, which provide rich background information and specific details for review generation.
[0117] The following is a specific example to introduce the comment prompt information in the embodiment of the present application.
[0118] Generate comment prompt information according to the structured template described above, which may include:
[0119] 1) Point of interest attribute information
[0120] Name: Oriental Pearl TV Tower
[0121] Nearby business district: Lujiazui, Pudong New Area, Shanghai
[0122] Distance from city center: 2.7 km
[0123] Theme style: leisure vacation
[0124] Crowd: Suitable for parents and children
[0125] Scene: Climbing high and looking far, tourist resorts, scenic spots, self-driving tours
[0126] Seasonal features: kite flying, climbing, and flower and leaf viewing
[0127] Attraction features: Glass plank road, sunrise viewing, night view
[0128] Certification: tourist areas, tourist attractions
[0129] Facilities: Exhibition Hall
[0130] Comments: Close to nature, popular attractions, beautiful night view, check-in place
[0131] Project: Wonderful performances, cycling
[0132] Play time: 2-3 hours
[0133] 2) Language style
[0134] Style: simple and clear style
[0135] Requirements: Use concise language to describe, avoid redundancy and complex sentences, and make the information clear and easy to understand.
[0136] 3) Theme Type
[0137] Theme: Scenery, History, Culture, Facilities
[0138] Description: For tourist attractions, focus on describing the beauty of the scenery, historical background, cultural characteristics, and convenience of facilities.
[0139] 4) Knowledge base information
[0140] There is no official parking lot at the destination. If you are driving, it is recommended to park in a nearby parking lot in advance. (Emotional tendency: neutral)
[0141] The room is located in the heart of downtown Shanghai, with a unique geographical location. (Emotional orientation: Positive)
[0142] Attractions include a fully transparent suspended observation corridor, a 230-story sky hotel, the Shanghai City History Development Exhibition Hall, a Coca-Cola Happy Restaurant, a space capsule, and a high-altitude VR roller coaster. (Emotional orientation: positive)
[0143] The 259-meter-long, fully transparent outdoor observation gallery offers excellent viewing opportunities, allowing users to take photos through the gap in the glass top using a selfie stick. (Emotional bias: Positive)
[0144] Business Hours: 12:00 PM - 10:00 PM daily. (Emotional Prediction: Neutral)
[0145] Address: Room 101, No. 17 Fuzhou Road. (Emotional orientation: Neutral)
[0146] You can enjoy the city scenery on both sides of the Huangpu River from different heights, especially from the fully transparent observation corridor at a height of 259 meters, which has a very good visual effect. (Emotional tendency: positive)
[0147] Getting there: Take the subway Line 2, exit 1, and take the escalator up to the circular island for photography. (Emotional bias: Neutral) The view faces the Oriental Pearl Tower, offering stunning night views, perfect for birthdays, dates, and anniversaries. (Emotional bias: Positive)
[0148] 5) Comment Template
[0149] Template structure:
[0150] Transportation: XXX.
[0151] Scenery: XXX.
[0152] Project recommendation: XXX.
[0153] Tickets: XXX
[0154] It should be noted that the structured template in the embodiment of the present application can be but is not limited to the above content, and can be tailored and expanded according to needs, and is not limited here.
[0155] The above step 204 , namely “inputting the comment prompt information and the image into a pre-trained comment generation model to generate candidate comments corresponding to the target interest point”, is described in detail below with reference to an embodiment.
[0156] In an embodiment of the present application, the comment generation model can be a multimodal large language model (MLLM), such as the open source multimodal large language models of Intern-VL and Qwen-VL. These models can simultaneously process multiple modal data such as text and images to generate high-quality candidate comments corresponding to the target interest points.
[0157] The comment prompt information and image are input into the multimodal large model. First, the comment prompt information and image need to be preprocessed, and the comment prompt information is converted into a text embedding vector, and the image is converted into an image embedding vector; then the text embedding vector and the image embedding vector are input into the pre-trained multimodal large model, and then the text embedding vector and the image embedding vector are fused through the multimodal fusion mechanism to obtain a multimodal feature representation; finally, based on the multimodal feature representation, candidate comments corresponding to the target interest point are generated.
[0158] Among them, the comment prompt information is converted into a text embedding vector that the model can understand. This is usually achieved through a pre-trained text encoder (such as a Transformer encoder) to map the text sequence to an embedding vector space of fixed dimension.
[0159] When converting an image into an image embedding vector, in an embodiment of the present application, the image uploaded by the user is first adjusted to a resolution that is suitable for the multimodal large model to obtain an adjusted target image, and then a pre-trained visual encoder (such as CLIP) is used to encode the target image to generate an image embedding vector.
[0160] This step is achieved by directly scaling the entire image, without segmenting it. This significantly reduces the number of image tokens, effectively lowering the computational complexity of large multimodal models. This improves processing speed and reduces resource consumption. This approach is particularly advantageous when processing multiple images, significantly improving system efficiency.
[0161] The text embedding vector and the image embedding vector are fused to obtain a multimodal feature representation, which is usually achieved through an attention mechanism or other fusion strategies, enabling the model to comprehensively consider text and image information.
[0162] Furthermore, after generating candidate comments corresponding to the target point of interest, the server sends the candidate comments to the user for the user to select and use.
[0163] In addition, in order to improve the reasoning efficiency of the multimodal large model, the embodiments of the present application also adopt a variety of optimization strategies:
[0164] Reasoning framework optimization: LMDeploy acceleration can be used. This framework provides more native support for Intern-VL and offers stronger accelerated reasoning performance than similar virtual large language models (vLLMs).
[0165] Model quantization technology: Using activation-aware weight quantization (AWQ) model quantization technology, it can reduce computational overhead while increasing processing speed, while generating almost no loss in candidate quality.
[0166] Specifically, the main principle of AWQ's model quantization technology is to first analyze the activations of a large multimodal model. Activations here refer to the outputs of neurons in each layer during the forward propagation process. Based on this activation information, the weights of the large multimodal model are divided into different groups. Within each group, the weights are balanced based on their relationship with the activation values, ensuring that the quantized weights retain the computational accuracy of the original model as much as possible. The grouped weights are then gradually quantized, typically using low-bit fixed-point representation (e.g., 8-bit or lower). At each quantization step, appropriate strategies, such as non-uniform quantization and dynamic range adjustment, are employed to maximize accuracy. Error feedback follows. The quantization process introduces certain errors. To minimize the impact of these errors on model performance, an error feedback mechanism can be introduced. After each quantization step, the quantization error is fed back to the next step for correction in subsequent quantization. Finally, a global optimization is performed on the entire large multimodal model, including weight adjustment and activation range reprocessing, to ensure balanced performance across different modalities.
[0167] Batch optimization: During the inference process, we use batch optimization strategies such as continuous batching and global batching to batch multiple comment generation requests and improve the efficiency of the graphics processing unit (GPU) on the server side.
[0168] Through the above optimization strategy, the multimodal large model can significantly reduce the use of video memory and improve inference performance while maintaining the same generation effect when converting comment prompt information and images into embedding vectors for fusion and generating candidate comments.
[0169] like Figure 3 As shown, the embodiment of the present application also provides a comment display method, which can be Figure 1 The method is executed by the user end of the system shown in FIG. The method includes the following steps:
[0170] Step 301: Displaying a comment publishing page for the target point of interest, the comment publishing page including an image upload area and a comment generation control;
[0171] Step 302: In response to the image upload area detecting that the user has uploaded at least one image and the comment generation control is triggered, the generated candidate comments are displayed. The candidate comments are generated based on the image and comment prompt information. The comment prompt information is generated based on the attribute information of the target interest point and the user's target emotional tendency towards the target interest point.
[0172] As can be seen, this embodiment of the application provides an image upload area and a comment generation control on the comment posting page. In response to the user's uploaded image and the comment generation control being triggered, candidate comments related to the target point of interest are generated. This method can effectively reduce the difficulty of user comment creation, improve the efficiency and quality of comment generation, and provide users with a more convenient interactive experience.
[0173] The steps in the above process and the effects that can be further produced are described in detail below with reference to embodiments.
[0174] The above step 301 , ie “displaying a comment publishing page for the target point of interest, wherein the comment publishing page includes an image upload area and a comment generation control”, is described in detail below in conjunction with an embodiment.
[0175] Typically, users can access the comment posting page for a specific POI by tapping it in the app. The comment posting page typically includes an image upload area and a comment input area. Users can upload relevant images in the image upload area, which typically uses intuitive icons and text prompts for user actions. It supports multiple image uploads and is compatible with various common image formats. Users can also enter their own comments in the comment input area.
[0176] In an embodiment of the present application, a comment generation control is also introduced on the comment posting page for the target point of interest. This control can be presented in the form of a button, icon, floating operation bar, or foldable panel. For example, when presented in the form of a button, the button can be labeled "Intelligently Generate Copy" or "One-Click to Get Comments". The comment generation control can be placed in a prominent position on the comment posting page, such as above the comment writing area or below the image upload area.
[0177] After a user has uploaded at least one image, if they find it difficult to write a review or want to get inspiration, they can click the review generation control. Or, if a user has uploaded at least one image and has written a partial review, and wants to further improve the review content and get more inspiration and suggestions, they can also click the review generation control.
[0178] The above step 302 , namely “displaying generated candidate comments in response to at least one image being uploaded to the image upload area and the comment generation control being triggered”, is described in detail below with reference to an embodiment.
[0179] In an embodiment of the present application, candidate comments are generated based on images uploaded by users and comment prompt information, wherein the comment prompt information is generated based on attribute information of the target point of interest and the user's target emotional tendency toward the target point of interest.
[0180] As an implementation, once a user successfully uploads at least one image in the image upload area and clicks the comment generation control, the system immediately responds to this sequence of actions. Based on the target POI's attribute information and the user's target sentiment toward the target POI, the server generates comment prompts. This comment prompt and at least one uploaded image are then fed into a pre-trained comment generation model to generate candidate comments corresponding to the target POI. The system then displays these candidate comments on the comment posting page.
[0181] As a feasible approach, a comment display area can be dynamically generated at a designated location on the comment publishing page, usually presented in the form of a concise and clear text box, to display the generated candidate comments to the user.
[0182] Furthermore, if the user enters comment information in the comment input area in addition to uploading the image, the target sentiment tendency is determined based on the comment information entered by the user. The comment information includes text comment information and / or rating information.
[0183] As another feasible method, when the user successfully uploads at least one image in the image upload area and enters comment information in the comment input area, clicking the comment generation control will dynamically generate a comment display area at a specified location on the comment publishing page, and display the generated candidate comments to the user.
[0184] That is, when the user successfully uploads at least one image in the image upload area, candidate comments can be generated and displayed to the user in the comment display area regardless of whether the user enters comment information.
[0185] In addition, in addition to the attribute information of the target interest point and the target emotional tendency of the user towards the target interest point, the comment prompt information based on which the candidate comments are generated can also be further determined by using the user's portrait information.
[0186] The process of generating candidate comments on the server side has been described in detail above and will not be repeated here.
[0187] Furthermore, after the candidate comments are displayed in the comment display area, a comment enable control and a comment discard control are also displayed in the comment display area. The comment enable control triggers the display of the candidate comments in the comment input area, allowing users to directly modify, supplement, or publish them, greatly simplifying the comment creation process. The comment discard control triggers the discarding of the candidate comments, allowing users to choose to regenerate the candidate comments or keep the current comment input area unchanged. This allows users to flexibly choose whether to use the system-generated candidate comments after uploading an image based on their needs and creation progress, improving the flexibility of comment creation and user experience.
[0188] As an example, Figure 4 As shown, this is the comment publishing page provided by the embodiment of this application. Figure 4 (a) The comment publishing page displays an image upload control 41, a comment generation control 42 (such as a "one-click comment generation" button), and a comment publishing control 43 (such as a "Publish" button). When a user clicks on a target point of interest in an application (such as a map application) and clicks the "Evaluate Now" button, the user enters the comment publishing page. By triggering the image upload control 41 on the comment publishing page, the user can upload at least one image related to the target point of interest. Then, by triggering the comment generation control 42 on the comment publishing page, the user sends a comment generation request to the server. The server generates candidate comments based on the above comment generation method and sends them to the user. Figure 4 (b) During the process of generating candidate comments on the server side, a window element can pop up on the comment publishing page. Before the candidate comments are generated, the window element can display a waiting message and the generation progress of the candidate comments, such as "Wait a moment, it will be ready soon (95%)"; Figure 4 (c) When the server generates candidate comments, the window element displays the generated candidate comments. At the same time, the window element also includes a comment use control 44 and a comment abandonment control 45. When the user triggers the comment use control 44, refer to Figure 4 (d) The generated candidate comments can be displayed in the comment input area. At this time, the user can modify and adjust them again, and finally publish the final comment content by triggering the comment publishing control 43. When the user triggers the comment abandonment control 45, refer to Figure 4 (a), the window element is closed and the page returns to Figure 4 (a) Status, or, reference Figure 4 (e) A pop-up window will pop up to ask the user whether to regenerate candidate comments. If the user clicks "Yes", the candidate comments will be regenerated for the user. If the user clicks "No", the page will return to the previous page. Figure 4 (a) Status.
[0189] The above diagram is only an example. In actual application, the page layout can be adjusted according to actual needs.
[0190] By introducing this intelligent comment generation function, the efficiency of user comment writing and the quality of comment generation can be greatly improved, thereby enhancing the user experience.
[0191] Based on the above review generation method, the embodiment of the present application also provides a training method for a review generation model, such as Figure 5 As shown, the following steps are included:
[0192] Step 501: Acquire training data including a plurality of training samples, wherein the training samples include at least one image sample uploaded by a user and a true value of a comment on a point of interest;
[0193] Step 502: Generate comment prompt information based on the attribute information of the point of interest and the target emotional tendency;
[0194] Step 503: Based on the comment prompt information and image samples, the comment generation model is trained, wherein the training includes: inputting the comment prompt information and image samples into the comment generation model to generate candidate comments corresponding to the points of interest; and updating the model parameters of the comment generation model using the loss function value corresponding to the training objective. The training objective includes: minimizing the difference between the candidate comments output by the comment generation model and the corresponding true values of the comments.
[0195] The training data consists of multiple training samples, each of which includes at least one image sample uploaded by a user and the true value of a comment about a point of interest. The true value of a comment about a point of interest is the actual comment about the point of interest.
[0196] To ensure the generality and robustness of the model, training data can be collected from multiple channels, covering different types of points of interest, diverse user groups, and various review styles.
[0197] In step 502, the implementation principle of generating comment prompt information is the same as the implementation principle of the above step 203, which will not be repeated here.
[0198] In step 503, the comment prompt information and the image sample are input into the comment generation model; the comment generation model converts the comment prompt information into a text embedding vector, converts the image sample into a target image with a resolution adapted to the comment generation model, converts the target image into an image embedding vector, and fuses the text embedding vector and the image embedding vector to obtain a multimodal feature representation; based on the multimodal feature representation, candidate comments corresponding to the point of interest are generated.
[0199] Furthermore, after the review generation model generates candidate reviews, it can also perform a quality assessment on them. Specifically, an automated evaluation tool is used to evaluate the candidate reviews, obtain evaluation results, and use the evaluation results as feedback to guide the update of model parameters and adjustment of training objectives of the review generation model.
[0200] Automated evaluation tools, such as advanced language models like GPT4o, can score candidate reviews across multiple dimensions, including but not limited to authenticity, diversity, content specificity, and image-text matching. Through a comprehensive evaluation of these dimensions, the quality of candidate reviews can be comprehensively measured.
[0201] For example, the prompt for automated evaluation could be: "You are a professional reviewer, and you need to make decisions based on factors such as high information content, close-to-realistic feel, and non-empty writing. Please judge which of the two copywritings is better based on the following information:
[0202] Copy 1: {man_out};
[0203] Copy 2:{model_out}.
[0204] Please directly tell me which copy is better, and keep your answer as concise as possible."
[0205] Alternatively, manual evaluation can be used. For example, by creating a web demo, users can compare their actual reviews with the candidate reviews generated by the review generation model from a real user perspective, assessing which reviews are more useful and provide more valuable information for users' decision-making. A metric for manual evaluation could be the win-loss ratio, which is the ratio of the model's results that are better than or equal to the manual results.
[0206] Through quality assessment, we can not only quickly iterate and optimize the review generation model, but also avoid falling into certain limitations, ensuring that the reviews generated by the model are more authentic, comprehensive and meet the actual needs of users.
[0207] like Figure 6 The figure shows a processing flow chart of the comment generation method provided by an embodiment of the present application. The server receives a comment generation request from a user, which includes identification information of the target point of interest, relevant images uploaded by the user, and the user's identity information. Based on the received identification information of the target point of interest, the attribute information of the target point of interest is determined, based on the user's identity information, the user's portrait information is determined, and at the same time, the user's target emotional tendency towards the target point of interest is determined, thereby generating a structured text, i.e., comment prompt information, based on the attribute information of the target point of interest, the user's target emotional tendency towards the target point of interest, and the user's portrait information.
[0208] The generated comment prompt information and the image uploaded by the user are then input into the multimodal large model (ie, the comment generation model in the embodiment of the present application) to generate candidate comments corresponding to the target point of interest.
[0209] Among them, this multimodal large model has adopted a number of optimization measures to improve reasoning efficiency, including reasoning framework optimization, model quantization technology, batch processing optimization, and image token optimization.
[0210] The generated review candidates undergo both automated and manual evaluation. Automated evaluation provides scoring across multiple dimensions, while manual evaluation displays the results via a web interface, allowing users or evaluators to compare and select. The evaluation results serve as feedback to guide model parameter updates and training objective adjustments, ultimately optimizing review quality.
[0211] The entire process emphasizes the dual mechanisms of multimodal information integration and quality assessment, aiming to improve the efficiency, quality and user experience of review generation.
[0212] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0213] According to another embodiment, a comment generating device is also provided. Figure 7 A schematic block diagram of a video generating apparatus according to an embodiment is shown. Figure 7 As shown, the apparatus 700 includes:
[0214] The request acquisition unit 701 is configured to acquire a comment generation request, wherein the comment generation request includes at least one image uploaded by a user;
[0215] An information determining unit 702 is configured to determine attribute information and target emotional tendency of the target point of interest;
[0216] A prompt generating unit 703 is configured to generate comment prompt information based on the attribute information of the target interest point and the target emotional tendency;
[0217] The comment generation unit 704 is configured to input the comment prompt information and the image into a pre-trained comment generation model to generate candidate comments corresponding to the target interest point.
[0218] Optionally, the comment generation request further includes: comment information of the user on the target point of interest, the comment information including text comment information and / or rating information;
[0219] The information determining unit 702 determines the target emotional tendency of the user towards the target point of interest and is specifically configured to:
[0220] Based on the user's comment information on the target point of interest, the target emotional tendency of the user towards the target point of interest is determined.
[0221] Optionally, the prompt generating unit 703 is configured to:
[0222] Filtering target knowledge items related to the attribute information of the target point of interest and in line with the target sentiment tendency from a pre-built retrieval-enhanced generation knowledge base;
[0223] Based on the attribute information of the target point of interest and the target knowledge item, the comment prompt information is generated according to a preset structured template.
[0224] Optionally, the search-enhanced knowledge base is constructed as follows:
[0225] Analyze historical review information of points of interest using a pre-trained text model to extract knowledge items that meet preset conditions. The knowledge items include the content of the knowledge points and the emotional tendency corresponding to the content of the knowledge points.
[0226] The knowledge items are stored in the search-enhanced generated knowledge base.
[0227] Optionally, the structured template includes point of interest attribute information and knowledge base information, and also includes at least one of a language style, a topic type, and a comment template.
[0228] Optionally, the prompt generating unit 703 is configured to generate the comment prompt information based on the attribute information of the target interest point, the target emotional tendency and the portrait information of the user.
[0229] Optionally, the comment generation unit 704 is configured to:
[0230] The comment prompt information and the image are input into the comment generation model; the comment generation model converts the comment prompt information into a text embedding vector, converts the image into a target image with a resolution adapted to the comment generation model, converts the target image into an image embedding vector, and fuses the text embedding vector and the image embedding vector to obtain a multimodal feature representation; based on the multimodal feature representation, a candidate comment corresponding to the target interest point is generated.
[0231] According to another embodiment, a comment display device is provided. Figure 8 A schematic block diagram of a video generating apparatus according to an embodiment is shown. Figure 8 As shown, the apparatus 800 includes:
[0232] A page display unit 801 is configured to display a comment publishing page of a target point of interest, wherein the comment publishing page includes an image upload area and a comment generation control;
[0233] The comment generation unit 802 is configured to display the generated candidate comments in response to the image upload area monitoring that the user has uploaded at least one image and the comment generation control is triggered. The candidate comments are generated based on the image and comment prompt information, and the comment prompt information is generated based on the attribute information of the target point of interest and the user's target emotional tendency towards the target point of interest.
[0234] Optionally, the comment publishing page further includes a comment input area;
[0235] The comment generation unit 802 is further configured to:
[0236] In response to at least one image being uploaded in the image upload area and the comment input area detecting comment information input by the user, the comment generation control is triggered to display the candidate comments, which are generated based on the image and the comment prompt information, and the comment prompt information is generated based on the attribute information of the target point of interest and the target emotional tendency, and the target emotional tendency is determined based on the comment information input by the user.
[0237] Optionally, the comment display area further includes a comment use control and a comment abandonment control, wherein the comment use control is used to trigger the display of the candidate comment in the comment input area; the comment abandonment control is used to trigger the abandonment of the candidate comment.
[0238] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.
[0239] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0240] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.
[0241] And an electronic device comprising:
[0242] one or more processors; and
[0243] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.
[0244] The present application also provides a computer program product, comprising a computer program, which implements the steps of any one of the methods described in the aforementioned method embodiments when executed by a processor.
[0245] in, Figure 9The electronic device architecture is shown as an example, and may include a processor 910, a video display adapter 911, a disk drive 912, an input / output interface 913, a network interface 914, and a memory 920. The processor 910, the video display adapter 911, the disk drive 912, the input / output interface 913, the network interface 914, and the memory 920 may be communicatively connected via a communication bus 930.
[0246] Among them, the processor 910 can be implemented by a general CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., to execute relevant programs to implement the technical solutions provided in this application.
[0247] The memory 920 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 920 can store an operating system 921 for controlling the operation of the electronic device 900, and a basic input and output system (BIOS) 922 for controlling the low-level operations of the electronic device 900. In addition, a web browser 923, a data storage management system 924, a comment generation device 700 / comment display device 800, etc. can also be stored. The above-mentioned comment generation device 700 / comment display device 800 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 920 and is called and executed by the processor 910.
[0248] The input / output interface 913 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0249] The network interface 914 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0250] The bus 930 comprises a pathway for transmitting information between the various components of the device (eg, the processor 910 , the video display adapter 911 , the disk drive 912 , the input / output interface 913 , the network interface 914 , and the memory 920 ).
[0251] It should be noted that although the above device only shows the processor 910, video display adapter 911, disk drive 912, input / output interface 913, network interface 914, memory 920, bus 930, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may also include only the components necessary to implement the solution of the present application, and does not necessarily include all the components shown in the figure.
[0252] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer program product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.
[0253] The above is a detailed introduction to the technical solutions provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the contents of this specification should not be understood as limiting this application.
Claims
1. A comment generation method, characterized in that: The method comprises: Obtaining a comment generation request for a target point of interest, the comment generation request including at least one image uploaded by a user; Determining attribute information of the target point of interest and the target emotional tendency of the user toward the target point of interest; Generate comment prompt information based on the attribute information of the target interest point and the target emotional tendency; The comment prompt information and the image are input into a pre-trained comment generation model to generate candidate comments corresponding to the target interest point.
2. The method according to claim 1, characterized in that The comment generation request further includes: comment information of the user on the target point of interest, wherein the comment information includes text comment information and / or rating information; Determining the target emotional tendency of the user towards the target point of interest includes: Based on the user's comment information on the target point of interest, the target emotional tendency of the user towards the target point of interest is determined.
3. The method according to claim 1, characterized in that The generating of comment prompt information based on the attribute information of the target interest point and the target emotional tendency includes: Filtering target knowledge items related to the attribute information of the target point of interest and in line with the target sentiment tendency from a pre-built retrieval-enhanced generation knowledge base; Based on the attribute information of the target point of interest and the target knowledge item, the comment prompt information is generated according to a preset structured template.
4. The method according to claim 3, characterized in that The search enhancement generation knowledge base is constructed as follows: Analyze historical review information of points of interest using a pre-trained text model to extract knowledge items that meet preset conditions. The knowledge items include the content of the knowledge points and the emotional tendency corresponding to the content of the knowledge points. The knowledge items are stored in the search-enhanced generated knowledge base.
5. The method according to claim 3, characterized in that The structured template includes point of interest attribute information and knowledge base information, and also includes at least one of a language style, a topic type, and a comment template.
6. The method according to claim 1, characterized in that The generating of comment prompt information based on the attribute information of the target interest point and the target emotional tendency includes: The comment prompt information is generated based on the attribute information of the target interest point, the target emotional tendency and the portrait information of the user.
7. The method according to any one of claims 1 to 6, characterized in that The step of inputting the comment prompt information and the image into a pre-trained comment generation model to generate candidate comments that match the comment sentiment tendency includes: The comment prompt information and the image are input into the comment generation model; the comment generation model converts the comment prompt information into a text embedding vector, converts the image into a target image adapted to the comment generation model, converts the target image into an image embedding vector, and fuses the text embedding vector and the image embedding vector to obtain a multimodal feature representation; and generates the candidate comment based on the multimodal feature representation.
8. A comment display method, characterized in that: The method comprises: Displaying a comment posting page for the target point of interest, wherein the comment posting page includes an image upload area and a comment generation control; In response to the image upload area detecting that the user has uploaded at least one image and the comment generation control being triggered, the generated candidate comments are displayed, the candidate comments are generated based on the image and comment prompt information, and the comment prompt information is generated based on the attribute information of the target point of interest and the user's target emotional tendency towards the target point of interest.
9. The method according to claim 8, characterized in that The comment posting page also includes a comment input area; In response to the image upload area detecting that the user has uploaded at least one image, and the comment generation control being triggered, displaying the generated candidate comments includes: In response to the image upload area detecting that the user has uploaded at least one image, and the comment input area detecting the comment information input by the user, the comment generation control is triggered and the generated candidate comments are displayed. The candidate comments are generated based on the image and the comment prompt information. The comment prompt information is generated based on the attribute information of the target interest point and the target emotional tendency. The target emotional tendency is determined based on the comment information input by the user.
10. A comment generation device, characterized in that: The device comprises: a request acquiring unit configured to acquire a comment generation request for a target point of interest, wherein the comment generation request includes at least one image uploaded by a user; an information determining unit, configured to determine attribute information and target emotional tendency of the target point of interest; a prompt generating unit configured to generate comment prompt information based on the attribute information and target emotional tendency of the target interest point; The comment generation unit is configured to input the comment prompt information and the image into a pre-trained comment generation model to generate candidate comments corresponding to the target interest point.
11. A comment display device, characterized in that: The device comprises: A page display unit is configured to display a comment publishing page of the target point of interest, wherein the comment publishing page includes an image upload area and a comment generation control; A comment generation unit is configured to display generated candidate comments in response to the image upload area monitoring that the user has uploaded at least one image and the comment generation control is triggered, wherein the candidate comments are generated based on the image and comment prompt information, and the comment prompt information is generated based on the attribute information of the target point of interest and the user's target emotional tendency towards the target point of interest.
12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.