Intelligent tour guide method, device, equipment and storage medium

CN116821492BActive Publication Date: 2026-08-18BEIJING XISOUND TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310763555.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-26
Publication Date
2026-08-18
Estimated Expiration
2043-06-26

AI Technical Summary

Technical Problem

[0004]然而,以上四种导览形式通常只提供固定信息,比如只提供固定的讲解内容和导览路线,缺乏智能性

Benefits of technology

[0085]本申请实施例提供的智能导览方案不但能够进行智能推荐,而且还能够进行智能语音导览。针对智能推荐,该方案会获取导览对象的兴趣数据和待游览区域的景点数据,进而基于获取到的兴趣数据和景点数据为导览对象生成推荐导览信息。由于在进行智能推荐时,该方案会对导览对象的兴趣爱好进行分析和理解,然后结合相关的景点数据为导览对象生成推荐导览信息,因此该方案实现了为不同导览对象进行个性化导览服务,较为智能化。针对智能语音导览,该方案不但能够与导览对象进行实时互动,而且在智能语音导览过程中还能够提供多种语言类型的导览服务,避免了导览对象在到达非母语地区旅游时可能遇到的语言障碍问题,该种导览方案的效果较佳,人机交互效率高。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116821492B_ABST
    Figure CN116821492B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent guide method and device, equipment and storage medium, and belongs to the technical field of Internet. The method comprises the following steps: obtaining interest data of a guide object and attraction data of a to-be-toured area; generating recommended guide information for the guide object based on the obtained interest data and attraction data; in the process of touring of the guide object based on the recommended guide information, determining a language type of voice data of the guide object in response to the obtained voice data; performing voice recognition on the voice data based on a voice recognition model matched with the language type; generating a response feedback matched with the recognized text, and determining a voice broadcast type of the response feedback; and broadcasting the response feedback by using a language variant matched with the determined voice broadcast type in the language type. The application can not only recommend personalized guide information for different guide objects, but also provide guide services in multiple language types in the process of intelligent voice guiding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technology, and in particular to an intelligent navigation method, device, equipment, and storage medium. Background Technology

[0002] The cultural tourism industry is closely related to people's leisure life, cultural activities, and experiential needs at both the material and spiritual levels. It is an economic form and industrial system led by the tourism, entertainment, service, and cultural industries. In recent years, with the vigorous development of the cultural tourism industry, many cities have made it an important part of industrial upgrading, and have begun to promote urban tourism upgrading from the perspective of all-for-one tourism.

[0003] Currently, guided tours primarily take four forms: human guides, audio guides, video guides, and written guides. Human guides are provided by professional guides or local guides. Audio guides use pre-recorded audio recordings, which tourists listen to via devices such as headphones. Video guides use pre-recorded videos, allowing tourists to learn about attractions and information. Written guides use text-based information, which tourists can access at the attraction or through printed materials.

[0004] However, the four types of guided tours mentioned above typically provide only fixed information, such as fixed explanations and routes, lacking any real-time interaction. Furthermore, guided tours are usually conducted in only one language, which can present language barriers for tourists visiting non-native language regions. For example, tourists may struggle to understand the guided tour content or miss important information. Summary of the Invention

[0005] This application provides an intelligent tour guide method, apparatus, device, and storage medium, which can not only recommend personalized tour guide information for different tour participants, but also provide tour guide services in multiple languages ​​during intelligent voice tours. The technical solution is as follows:

[0006] On the one hand, an intelligent navigation method is provided, the method comprising:

[0007] Acquire the interest data of the guided tourists and the attraction data of the areas to be visited;

[0008] Based on the interest data and the attraction data, recommended tour information is generated for the tour target.

[0009] During the tour of the guided object based on the recommended tour information, in response to obtaining the voice data of the guided object, the language type of the voice data is determined;

[0010] Based on a speech recognition model that matches the language type, speech recognition is performed on the speech data;

[0011] Generate a response that matches the identified text, and determine the voice broadcast type of the response;

[0012] The response feedback is broadcast using a language variant that matches the voice broadcast type under the specified language type.

[0013] In one possible implementation, determining the voice broadcast type of the response feedback includes:

[0014] Obtain the current physiological health status and emotional state of the guided object;

[0015] Obtain the gender and age attributes of the navigation object;

[0016] Based on the gender, age, current physical health, and emotional state of the guided tour participants, the type of voice broadcast for the response feedback is determined.

[0017] In one possible implementation, determining the voice broadcast type of the response feedback based on the gender attribute, age attribute, current physical health status, and emotional state of the tour guide includes:

[0018] Based on the gender and age attributes of the guided tour object, the first voice broadcast type of the response feedback is determined;

[0019] Based on the current physiological health and emotional state of the guided object, determine the second voice broadcast type of the response feedback;

[0020] The voice broadcast type that matches the interest data of the guided object in the first and second voice broadcast types is determined as the final voice broadcast type for the response feedback.

[0021] In one possible implementation, determining the language type of the voice data includes:

[0022] Acoustic feature extraction is performed on the speech data to obtain the filter bank features and speaker pitch features of the speech data;

[0023] The discrete cosine transform of the filter bank features is used to obtain the Mel-frequency cepstral coefficients of the speech data.

[0024] The speaker's pitch features and the Mel cepstral coefficients are fused to obtain the fused acoustic features of the speech data.

[0025] Based on the fusion acoustic features and language type recognition model, the speech data is used to identify the language type and obtain the language type adopted by the speech data.

[0026] In one possible implementation, the step of performing language type identification on the speech data based on the fused acoustic features and language type recognition model to obtain the language type adopted by the speech data includes:

[0027] The fused acoustic features are input into a target neural network; wherein the target neural network is a natural language processing model including an encoder and a decoder;

[0028] The output features of the encoder are input into the language type recognition model;

[0029] Based on the output of the language type recognition model, the language type of the speech data is determined.

[0030] In one possible implementation, the training process of the speech recognition model matching the language type includes:

[0031] For each of the multiple language types, a training sample set for that language type is obtained; the training sample set includes a first type of sample speech using that language type collected from the network.

[0032] Based on the first type of sample speech, the training sample set is expanded to obtain the second type of sample speech using the language type.

[0033] Based on the first type of sample speech and the second type of sample speech, a speech recognition model matching the language type is trained.

[0034] In one possible implementation, obtaining the interest data of the guided object and the attraction data of the area to be visited includes:

[0035] The question types of the historical voice data input by the tour guide and the tour guide's evaluation of the historically recommended tour information are obtained to obtain the tour guide's interest data.

[0036] The popularity, ratings, and attributes of attractions in the area to be visited are obtained to acquire attraction data for the area.

[0037] In one possible implementation, the step of outputting recommended tour information for the tour target based on the interest data and the attraction data includes:

[0038] After preprocessing the data related to the tour object, feature encoding is performed on the preprocessed data of the tour object to obtain the interest features of the tour object;

[0039] After preprocessing the data related to the area to be visited, feature encoding is performed on the preprocessed data of the area to be visited to obtain the scenic spot features of the area to be visited.

[0040] The interest characteristics of the guided object and the scenic spot characteristics of the area to be visited are fused together;

[0041] The fused features are input into the recommendation model to obtain the initial recommended tour information for the tour object; wherein, the initial recommended tour information includes multiple scenic spot tour routes arranged in sequence;

[0042] Obtain the gender, age, occupation type, current physical health status, and emotional state of the guided tour participants;

[0043] Based on at least one of the physiological health status, emotional state, gender attribute, age attribute, and occupation type, the ranking results of the multiple scenic spot tour routes are readjusted to obtain the final recommended tour information for the tour target.

[0044] On the other hand, a smart tour guide device is provided, the device comprising:

[0045] The acquisition module is configured to acquire the interest data of the guided users and the attraction data of the area to be visited.

[0046] The first generation module is configured to output recommended tour information for the tour target based on the interest data and the attraction data;

[0047] The determination module is configured to, in response to acquiring the voice data of the guided object during the tour based on the recommended guided information, determine the language type of the voice data.

[0048] The speech recognition module is configured to perform speech recognition on the speech data based on a speech recognition model that matches the language type.

[0049] The second generation module is configured to generate response feedback that matches the recognized text;

[0050] The output module is configured to determine the voice broadcast type of the response feedback; and to broadcast the response feedback using a language variant that matches the voice broadcast type under the language type.

[0051] In one possible implementation, the output module is configured as follows:

[0052] Obtain the current physiological health status and emotional state of the guided object;

[0053] Obtain the gender and age attributes of the navigation object;

[0054] Based on the gender, age, current physical health, and emotional state of the guided tour participants, the type of voice broadcast for the response feedback is determined.

[0055] In one possible implementation, the output module is configured as follows:

[0056] Based on the gender and age attributes of the guided tour object, the first voice broadcast type of the response feedback is determined;

[0057] Based on the current physiological health and emotional state of the guided object, determine the second voice broadcast type of the response feedback;

[0058] The voice broadcast type that matches the interest data of the guided object in the first and second voice broadcast types is determined as the final voice broadcast type for the response feedback.

[0059] In one possible implementation, the determining module is configured as follows:

[0060] Acoustic feature extraction is performed on the speech data to obtain the filter bank features and speaker pitch features of the speech data;

[0061] The discrete cosine transform of the filter bank features is used to obtain the Mel-frequency cepstral coefficients of the speech data.

[0062] The speaker's pitch features and the Mel cepstral coefficients are fused to obtain the fused acoustic features of the speech data.

[0063] Based on the fusion acoustic features and language type recognition model, the speech data is used to identify the language type and obtain the language type adopted by the speech data.

[0064] In one possible implementation, the determining module is configured as follows:

[0065] The fused acoustic features are input into a target neural network; wherein the target neural network is a natural language processing model including an encoder and a decoder;

[0066] The output features of the encoder are input into the language type recognition model;

[0067] Based on the output of the language type recognition model, the language type of the speech data is determined.

[0068] In one possible implementation, the training process of the speech recognition model matching the language type includes:

[0069] For each of the multiple language types, a training sample set for that language type is obtained; the training sample set includes a first type of sample speech using that language type collected from the network.

[0070] Based on the first type of sample speech, the training sample set is expanded to obtain the second type of sample speech using the language type.

[0071] Based on the first type of sample speech and the second type of sample speech, a speech recognition model matching the language type is trained.

[0072] In one possible implementation, the acquisition module is configured as follows:

[0073] The question types of the historical voice data input by the tour guide and the tour guide's evaluation of the historically recommended tour information are obtained to obtain the tour guide's interest data.

[0074] The popularity, ratings, and attributes of attractions in the area to be visited are obtained to acquire attraction data for the area.

[0075] In one possible implementation, the first generation module is configured as follows:

[0076] After preprocessing the data related to the tour object, feature encoding is performed on the preprocessed data of the tour object to obtain the interest features of the tour object;

[0077] After preprocessing the data related to the area to be visited, feature encoding is performed on the preprocessed data of the area to be visited to obtain the scenic spot features of the area to be visited.

[0078] The interest characteristics of the guided object and the scenic spot characteristics of the area to be visited are fused together;

[0079] The fused features are input into the recommendation model to obtain the initial recommended tour information for the tour object; wherein, the initial recommended tour information includes multiple scenic spot tour routes arranged in sequence;

[0080] Obtain the gender, age, occupation type, current physical health status, and emotional state of the guided tour participants;

[0081] Based on at least one of the physiological health status, emotional state, gender attribute, age attribute, and occupation type, the ranking results of the multiple scenic spot tour routes are readjusted to obtain the final recommended tour information for the tour target.

[0082] On the other hand, a computer device is provided, the device including a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to implement the above-described intelligent navigation method.

[0083] On the other hand, a computer-readable storage medium is provided, wherein at least one piece of program code is stored in the storage medium, the at least one piece of program code being loaded and executed by a processor to implement the above-described intelligent navigation method.

[0084] On the other hand, a computer program product or computer program is provided, which includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform the above-described intelligent navigation method.

[0085] The intelligent tour guide solution provided in this application not only enables intelligent recommendations but also intelligent voice guidance. For intelligent recommendations, the solution acquires the interests of the tour participant and the attraction data of the area to be visited, and then generates recommended tour information based on this data. Because the solution analyzes and understands the tour participant's interests and then combines this with relevant attraction data to generate recommended tour information, it achieves personalized tour services for different tour participants, making it quite intelligent. For intelligent voice guidance, the solution not only allows real-time interaction with the tour participant but also provides multilingual guidance services during the intelligent voice tour, avoiding language barriers that tour participants may encounter when traveling to non-native language areas. This tour guide solution is highly effective and has high human-computer interaction efficiency. Attached Figure Description

[0086] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0087] Figure 1 This is a schematic diagram of the implementation environment of an intelligent navigation method provided in an embodiment of this application;

[0088] Figure 2 This is a flowchart of an intelligent navigation method provided in an embodiment of this application;

[0089] Figure 3This is a flowchart of another intelligent navigation method provided in an embodiment of this application;

[0090] Figure 4 This is a schematic diagram of the structure of an intelligent tour guide device provided in an embodiment of this application;

[0091] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application;

[0092] Figure 6 This is a schematic diagram of the structure of another computer device provided in an embodiment of this application. Detailed Implementation

[0093] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0094] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items that have essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or execution order. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms.

[0095] These terms are simply used to distinguish one element from another. For example, without departing from the various examples, the first element can be referred to as the second element, and similarly, the second element can be referred to as the first element. Both the first and second elements can be elements, and in some cases, they can be separate and distinct elements.

[0096] "At least one" refers to one or more elements. For example, at least one element can be one element, two elements, three elements, or any integer number of elements greater than or equal to one. "Multiple" refers to two or more elements. For example, multiple elements can be two elements, three elements, or any integer number of elements greater than or equal to two.

[0097] In this article, "and / or" indicates that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0098] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0099] As is well known, guided tours play a vital role in cultural and tourism services. On the one hand, they provide tourists with accurate, comprehensive, and engaging explanations of attractions and historical and cultural introductions, helping them better understand the background and characteristics of their destinations and enhancing their cultural literacy and travel experience. On the other hand, guided tours also offer tourist attractions an important service method and value-added service, improving their competitiveness and attractiveness. Through guided tours, tourism managers can provide tourists with more professional and high-quality services, increasing tourist satisfaction and loyalty, and boosting return visits and positive word-of-mouth. However, current guided tour services suffer from low levels of intelligence, resulting in a generally poor tourist experience. The current guided tour solutions mainly have the following shortcomings:

[0100] 1. Lack of personalization: Current cultural and tourism guided tours typically only offer fixed routes and the same information, such as fixed explanations and tour routes, and cannot be adjusted or customized according to tourists' individual needs and interests. This may lead to tourists feeling bored or dissatisfied.

[0101] 2. Technical limitations: Current cultural and tourism guides usually use human voice narration or guide text, which may be subject to technical limitations, such as poor sound quality or difficulty in reading the text.

[0102] 3. Insufficient interactivity: Current cultural and tourism guided tours often lack interactivity, failing to engage in real-time interaction and communication with tourists. This may make tourists feel lonely or lack a sense of participation.

[0103] 4. Lack of innovation: The current way of guiding cultural and tourism tours has been around for a long time, and tourists may have grown tired of it. Therefore, more innovative ways are needed to attract their attention.

[0104] 5. Language Barriers: Tourists may encounter language barriers, especially when traveling to areas where the language is not their native tongue. Current tourism guides typically only offer audio narration in one language, which may make it difficult for tourists to understand or cause them to miss important information.

[0105] To address the shortcomings of the aforementioned technical solutions, this application proposes an intelligent tour guide solution. This solution combines a deep learning model (such as the currently advanced language model ChatGPT) and trains it with massive amounts of corpus data to achieve powerful natural language understanding and generation capabilities. It can provide users with personalized travel guides, support user interaction, and offer tour guide services in multiple languages, thus solving the language barriers that may be encountered when traveling in non-native language regions.

[0106] The intelligent navigation solution provided in this application will be described below through the following implementation methods.

[0107] Figure 1 This is a schematic diagram of the implementation environment of an intelligent navigation method provided in this application embodiment.

[0108] In this embodiment of the application, the implementation environment includes a computer device. For example, see [link to relevant documentation]. Figure 1 The aforementioned computer equipment includes a terminal 101 and a server 102; in other words, the intelligent navigation method is jointly executed by the terminal 101 and the server 102. Alternatively, the intelligent navigation method can also be executed by the terminal 101 alone, and this application does not limit this to that.

[0109] For example, terminal 101 is a computer device with a display screen, such as a smartphone or tablet computer; while server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and this application does not limit it.

[0110] In addition, the server involved in the embodiments of this application may also include other servers in order to provide more comprehensive and diversified services.

[0111] Furthermore, those skilled in the art will understand that the number of terminals may be more or less than shown in the figures. For example, the number of terminals may be only a few, or dozens or hundreds, or even more; this application does not limit the number.

[0112] For example, an application providing cultural and tourism services is installed on terminal 101. For instance, tourists can download the application via their mobile phones or tablets and then access cultural and tourism services through the application. Server 102 provides background services for the application. The application is also referred to as a tour guide application, and tourists are referred to as tour guide objects or users in this embodiment.

[0113] Figure 2 This is a flowchart illustrating an intelligent navigation method provided in an embodiment of this application. The method is executed by a computer device, such as a terminal and a server working together. See also... Figure 2The method flow provided in this application embodiment includes:

[0114] 201. The computer equipment acquires the interest data of the tour participants and the attraction data of the area to be visited; based on the interest data and attraction data of the tour participants, it generates recommended tour information for the tour participants.

[0115] In this embodiment, the intelligent navigation service mainly includes two aspects: intelligent recommendation and intelligent voice navigation. This step is used to implement intelligent recommendation, which will be described below.

[0116] This part uses deep learning models (such as the advanced language model ChatGPT) to analyze and understand tourists' interests and then combine relevant attraction data to generate recommended tour information for tourists.

[0117] For example, the area to be visited is a geographical region that includes at least one attraction, such as a certain urban area, a certain scenic area, or a certain geographical range; this application does not limit this. The aforementioned attraction data includes, but is not limited to, the popularity, ratings, and attributes of attractions in the area to be visited. The aforementioned recommended tour information includes, but is not limited to, tour routes that best suit tourists' interests and introductions to the attractions.

[0118] In one possible implementation, attraction popularity could be the frequency of online searches for the attraction's keywords, the attraction's visitor traffic during a certain period, or the number of tickets sold during a certain period; this application does not limit this. Furthermore, the aforementioned attraction attributes include, but are not limited to, the attraction's name, features, province, attraction level, address, latitude and longitude, brief description, opening hours, suitable season, or recommended visit time; this application does not limit this.

[0119] In another possible implementation, intelligent recommendation mainly includes the following steps:

[0120] 1. Collect data related to tourists. For example, obtain the types of questions users have asked in their history and their evaluations of each recommended tour information provided.

[0121] 2. Collect data related to the area to be visited, such as data on attractions within the scenic area.

[0122] 3. Data preprocessing. The collected data is cleaned and processed, such as data deduplication, missing value imputation, and correction. This application does not limit the scope of this process.

[0123] 4. Train a deep learning model. For example, use deep learning techniques to train a deep learning model to analyze and understand tourists' interests and preferences.

[0124] 5. Process the tourist attraction data. For example, use deep learning technology to train another deep learning model to analyze and process the tourist attraction data.

[0125] 6. Recommendation Algorithm. Combining the analysis and understanding of tourists' interests and preferences with the analysis and processing of attraction data, a recommendation algorithm is used to provide personalized tour recommendations for tourists, resulting in recommended outcomes.

[0126] 7. Results Output: Display recommended results to tourists, such as attraction introductions and tour route planning.

[0127] 202. During the tour of the guided object based on the recommended guided information, in response to the acquisition of the guided object's voice data, the computer device determines the language type of the acquired voice data; based on the voice recognition model matching the language type, it performs voice recognition on the acquired voice data.

[0128] This step and step 203 below are used to implement intelligent voice guidance, which is described below. This part uses speech recognition technology to convert the tourist's voice input into text (supporting multiple language types), then uses natural language processing technology for semantic understanding and response generation, and finally uses speech synthesis technology to output the response as audio for the tourist to hear.

[0129] 203. The computer device generates and recognizes the text-matched response feedback, and determines the voice broadcast type of the response feedback; it then broadcasts the response feedback using a language variant that matches the voice broadcast type under that language type.

[0130] Language variants refer to different forms of expression within a single language type. Taking Chinese as an example, Chinese includes Standard Mandarin and various regional dialects, both of which are language variants of Chinese.

[0131] In one possible implementation, intelligent voice navigation mainly includes the following steps:

[0132] 1. Collect voice data: Collect voice data of questions that tourists may ask. This voice data is used to train the voice recognition model.

[0133] 2. Train a speech recognition model: Use deep learning technology to train a speech recognition model to convert speech input into text.

[0134] 3. Integrate ChatGPT: Integrate ChatGPT into the navigation application to generate answers.

[0135] 4. Dialogue Generation: Semantic understanding of tourists' voice input is performed using natural language processing technology, and then ChatGPT is used to generate response text that matches the voice input.

[0136] 5. Speech Synthesis: Convert the answer text into speech and output it to the tourists.

[0137] The intelligent tour guide solution provided in this application not only enables intelligent recommendations but also intelligent voice guidance. For intelligent recommendations, the solution acquires the interests of the tour participant and the attraction data of the area to be visited, and then generates recommended tour information based on this data. Because the solution analyzes and understands the tour participant's interests and then combines this with relevant attraction data to generate recommended tour information, it achieves personalized tour services for different tour participants, making it quite intelligent. For intelligent voice guidance, the solution not only allows real-time interaction with the tour participant but also provides multilingual guidance services during the intelligent voice tour, avoiding language barriers that tour participants may encounter when traveling to non-native language areas. This tour guide solution is highly effective and has high human-computer interaction efficiency.

[0138] The above is only a brief introduction to some technical details of the recommended solution provided in the embodiments of this application. The following is based on... Figure 3 The illustrated embodiments provide a detailed description of the recommended solution.

[0139] Figure 3 This is a flowchart of another intelligent navigation method provided in an embodiment of this application. The method is executed by a computer device, such as a terminal and a server working together. See also... Figure 3 The method flow provided in this application embodiment includes:

[0140] 301. Computer equipment acquires interest data of the guided visitors and attraction data of the areas to be visited.

[0141] For example, the interest data of the tour guide and the attraction data of the area to be visited are obtained, including but not limited to: obtaining the question types of the historical voice data input by the tour guide and the tour guide's evaluation of the historically recommended tour information to obtain the tour guide's interest data; obtaining the attraction popularity, attraction evaluation and attraction attributes of the area to be visited to obtain the attraction data of the area to be visited.

[0142] It should be noted that, in addition to the examples mentioned above, other types of data may be included for the interest data of the guided tour participants and the attraction data of the area to be visited. This application does not limit this type of data.

[0143] 302. The computer equipment generates recommended tour information for the tour participants based on their interest data and attraction data.

[0144] In one possible implementation, recommended tour information is output to the tour user based on the user's interest data and the attraction data of the area to be visited, including but not limited to the following methods:

[0145] 3021. After preprocessing the data related to the tour object, feature encoding is performed on the preprocessed tour object data to obtain the tour object's interest features.

[0146] For example, this step involves mapping the interest data of the guided tour object to a feature space, thereby obtaining the interest features of the guided tour object. This can be achieved by training an interest analysis model based on deep learning technology, and then using this interest analysis model to perform feature encoding on the preprocessed guided tour object data to obtain the interest features of the guided tour object; this application does not limit this approach.

[0147] 3022. After preprocessing the data related to the area to be visited, feature encoding is performed on the preprocessed data of the area to be visited to obtain the scenic spot features of the area to be visited.

[0148] For example, this step involves mapping the attraction data of the area to be visited to a feature space, thereby obtaining the attraction features of the area to be visited. This can be achieved by training an attraction analysis model based on deep learning technology, and then using this model to encode the features of the preprocessed data of the area to be visited, thus obtaining the attraction features of the area to be visited. This application does not limit this approach.

[0149] 3023. The interest characteristics of the tour participants and the scenic spot characteristics of the area to be visited are fused together; the fused features are input into the recommendation model to obtain the initial recommended tour information for the tour participants; wherein, the initial recommended tour information includes multiple scenic spot tour routes arranged in sequence.

[0150] For example, the above feature fusion step may involve concatenating the interest features of the guided object and the attraction features of the area to be visited; this application is not limited to this. Alternatively, a fusion layer may be used to fuse the interest features of the guided object and the attraction features of the area to be visited. The fusion layer can be either a linear mapping layer or a column exchange layer; this application is not limited to this. Furthermore, the parameters of the fusion layer can be randomly initialized and optimized based on a backpropagation algorithm, such as a stochastic gradient descent algorithm.

[0151] In addition, if a feature has a relatively high dimensionality, it can be reduced in dimensionality first, and then feature fusion can be performed.

[0152] In addition to the sightseeing route, the initial recommended tour information may also include introductions to the attractions, but this application does not impose any restrictions on this.

[0153] In another possible implementation, the above-mentioned interest analysis model, the above-mentioned scenic spot analysis model, and the above-mentioned recommendation model can be trained in the following way:

[0154] A training sample set is obtained, which includes massive amounts of tourist interest data, massive amounts of attraction data for tourist areas, and historical visit data of these tourists in the aforementioned tourist areas. Based on this training sample set, a deep learning technique is used to train an interest analysis model, a attraction analysis model, and a recommendation model until the models converge. For example, the model convergence condition can be that the loss value calculated by the loss function is less than a preset error value. The outputs of the interest analysis model and the attraction analysis model serve as the inputs to the recommendation model; the aforementioned loss function is a loss function constructed for the recommendation model, and the aforementioned loss value is calculated based on the aforementioned historical visit data and the predicted tour information output by the recommendation model; this application does not limit this aspect.

[0155] 3024. Obtain the gender, age, occupation, current physical health status, and emotional state of the tour participants.

[0156] It should be noted that the gender, age, occupation, current physical health and emotional state of the above-mentioned tour participants are all authorized by the users or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0157] 3025. Based on at least one of the following: physiological health status, emotional state, gender attribute, age attribute, and occupation type, readjust the ranking results of multiple scenic spot tour routes to obtain the final recommended tour information for the tour participants.

[0158] For example, if the tour guide's emotional state is negative, a quiet and peaceful tour route will be recommended first; if the tour guide's physical health is recovering from a serious illness, a shorter tour route with rest areas will be recommended first; if the tour guide's gender is male, a tour route with fitness facilities or exciting activities will be recommended first.

[0159] 303. During the tour of the guided object based on the recommended guided information, in response to the acquisition of the guided object's voice data, the computer device determines the language type of the acquired voice data.

[0160] In one possible implementation, the embodiments of this application determine the language type used for the voice input of the tour guide object in the following manner:

[0161] 3031. Perform acoustic feature extraction on the speech data of the guided object to obtain the filter bank features and speaker pitch features of the speech data.

[0162] Among them, filter bank features refer to FilterBank features, also known as FBank features, and speaker pitch features refer to pitch features.

[0163] 3032. Perform discrete cosine transform on the filter bank features to obtain the Mel cepstral coefficients of the speech data.

[0164] Among them, the Mel frequency cepstrum coefficient is a characteristic of MFCC (Mel Frequency Cepstrum Coefficient).

[0165] 3033. The speaker's pitch features and Mel cepstral coefficients are fused to obtain the fused acoustic features of the speech data.

[0166] For example, speaker pitch features and Mel-frequency cepstral coefficients can be input into a fusion layer for feature fusion. This fusion layer can be either a linear mapping layer or a column exchange layer; this application does not limit its application to either. Furthermore, the parameters of the fusion layer can be randomly initialized and optimized using a backpropagation algorithm, such as a stochastic gradient descent algorithm.

[0167] 3034. Based on the fusion of acoustic features and language type recognition model, the language type of the speech data is identified to obtain the language type used in the speech data.

[0168] In another possible implementation, based on a fusion of acoustic features and a language type recognition model, language type identification is performed on the speech data to obtain the language type adopted by the speech data, including but not limited to:

[0169] The fused acoustic features are input into the target neural network, which is a natural language processing model including an encoder and a decoder. The output features of the encoder are input into the language type recognition model. Based on the output of the language type recognition model, the language type of the speech data of the tour guide object is determined.

[0170] 304. The computer device performs speech recognition on the acquired speech data based on a speech recognition model that matches the language type.

[0171] In one possible implementation, the training process of the speech recognition model matching the aforementioned language types includes, but is not limited to:

[0172] For each of the multiple language types, a training sample set for the aforementioned language type is obtained; the training sample set includes a first type of sample speech collected from the network that adopts the aforementioned language type; based on the first type of sample speech, the training sample set is expanded to obtain a second type of sample speech that adopts the aforementioned language type; based on the first type of sample speech and the second type of sample speech, a speech recognition model matching the aforementioned language type is trained.

[0173] The aforementioned language types include, but are not limited to, Chinese, English, Russian, Arabic, French, and Spanish, etc., and this application does not limit them.

[0174] For example, text expansion can be performed by replacing keywords that appear in the recognized text of the first type of sample speech, and this application does not limit this to either.

[0175] 305. The computer device generates and recognizes a response feedback that matches the text, and determines the voice broadcast type of the response feedback.

[0176] The types of voice broadcasts include, but are not limited to: ordinary male voice, ordinary female voice, elderly voice, child voice, loli voice, celebrity voice, cheerful voice, or broadcasting voice, etc. This application does not limit these types of voice broadcasts.

[0177] In one possible implementation, the type of voice broadcast for the response feedback is determined, including but not limited to:

[0178] 3051. Obtain the current physiological health status, emotional state, gender attribute, and age attribute of the guided object.

[0179] It should be noted that the physiological health status, emotional state, gender attributes, and age attributes of the aforementioned guided individuals are all authorized by the users or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0180] 3052. Based on the gender, age, current physical health and emotional state of the tour guide, determine the type of voice broadcast for the response feedback.

[0181] In another possible implementation, the type of voice broadcast for the response feedback is determined based on the gender, age, current physical health, and emotional state of the tour guide, including but not limited to:

[0182] Based on the gender and age attributes of the tour participants, the first voice broadcast type for the response feedback is determined; for example, older men might prefer a broadcast-style voice. Based on the current physical and emotional state of the tour participants, the second voice broadcast type for the response feedback is determined; assuming the tour participants are currently healthy and in a good mood, they might prefer a cheerful tone. The voice broadcast type that matches the tour participants' interest data between the first and second voice broadcast types is determined as the final voice broadcast type for the response feedback. For example, if the voice broadcast type matching the tour participants' interest data is a broadcast-style voice, then the final voice broadcast type for the response feedback will be a broadcast-style voice.

[0183] 306. The computer equipment uses a language variant that matches the voice broadcast type under this language type to broadcast response feedback.

[0184] This step is by Figure 1 The terminal shown is used to broadcast response feedback for the navigation object.

[0185] The intelligent tour guide solution provided in this application not only enables intelligent recommendations but also intelligent voice guidance. For intelligent recommendations, the solution acquires the interests of the tour participant and the attraction data of the area to be visited, and then generates recommended tour information based on this data. Because the solution analyzes and understands the tour participant's interests and then combines this with relevant attraction data to generate recommended tour information, it achieves personalized tour services for different tour participants, making it quite intelligent. For intelligent voice guidance, the solution not only allows real-time interaction with the tour participant but also provides multilingual guidance services during the intelligent voice tour, avoiding language barriers that tour participants may encounter when traveling to non-native language areas. This tour guide solution is highly effective and has high human-computer interaction efficiency.

[0186] For example, by combining current deep learning models (such as the advanced language model ChatGPT), this application embodiment can provide tourists with intelligent and personalized guided tour services, and support interaction with tourists, greatly enhancing the enjoyment. Furthermore, this application embodiment can continue to train the recommendation model based on user evaluations of the recommended guided tour information, thereby continuously optimizing the guided tour service.

[0187] In addition, the embodiments of this application solve the problem that related Chinese travel guide services can only provide fixed information and cannot adjust and customize the guide service according to the individual needs and interests of tourists.

[0188] In summary, the embodiments of this application enable real-time interaction and communication with tourists, providing interactive, personalized, and engaging intelligent service experiences for tourists during cultural and tourism guided tours.

[0189] Figure 4 This is a schematic diagram of the structure of an intelligent tour guide device provided in an embodiment of this application. See also... Figure 4 The device includes:

[0190] The acquisition module 401 is configured to acquire the interest data of the guided object and the attraction data of the area to be visited.

[0191] The first generation module 402 is configured to output recommended tour information for the tour object based on the interest data and the attraction data;

[0192] The determination module 403 is configured to, in response to acquiring the voice data of the guided object during the guided object's tour based on the recommended guided information, determine the language type used in the voice data;

[0193] The speech recognition module 404 is configured to perform speech recognition on the speech data based on a speech recognition model that matches the language type.

[0194] The second generation module 405 is configured to generate response feedback that matches the recognized text;

[0195] Output module 406 is configured to determine the voice broadcast type of the response feedback; and broadcast the response feedback using a language variant that matches the voice broadcast type under the language type.

[0196] The intelligent tour guide solution provided in this application not only enables intelligent recommendations but also intelligent voice guidance. For intelligent recommendations, the solution acquires the interests of the tour participant and the attraction data of the area to be visited, and then generates recommended tour information based on this data. Because the solution analyzes and understands the tour participant's interests and then combines this with relevant attraction data to generate recommended tour information, it achieves personalized tour services for different tour participants, making it quite intelligent. For intelligent voice guidance, the solution not only allows real-time interaction with the tour participant but also provides multilingual guidance services during the intelligent voice tour, avoiding language barriers that tour participants may encounter when traveling to non-native language areas. This tour guide solution is highly effective and has high human-computer interaction efficiency.

[0197] In one possible implementation, the output module is configured as follows:

[0198] Obtain the current physiological health status and emotional state of the guided object;

[0199] Obtain the gender and age attributes of the navigation object;

[0200] Based on the gender, age, current physical health, and emotional state of the guided tour participants, the type of voice broadcast for the response feedback is determined.

[0201] In one possible implementation, the output module is configured as follows:

[0202] Based on the gender and age attributes of the guided tour object, the first voice broadcast type of the response feedback is determined;

[0203] Based on the current physiological health and emotional state of the guided object, determine the second voice broadcast type of the response feedback;

[0204] The voice broadcast type that matches the interest data of the guided object in the first and second voice broadcast types is determined as the final voice broadcast type for the response feedback.

[0205] In one possible implementation, the determining module is configured as follows:

[0206] Acoustic feature extraction is performed on the speech data to obtain the filter bank features and speaker pitch features of the speech data;

[0207] The discrete cosine transform of the filter bank features is used to obtain the Mel-frequency cepstral coefficients of the speech data.

[0208] The speaker's pitch features and the Mel cepstral coefficients are fused to obtain the fused acoustic features of the speech data.

[0209] Based on the fusion acoustic features and language type recognition model, the speech data is used to identify the language type and obtain the language type adopted by the speech data.

[0210] In one possible implementation, the determining module is configured as follows:

[0211] The fused acoustic features are input into a target neural network; wherein the target neural network is a natural language processing model including an encoder and a decoder;

[0212] The output features of the encoder are input into the language type recognition model;

[0213] Based on the output of the language type recognition model, the language type of the speech data is determined.

[0214] In one possible implementation, the training process of the speech recognition model matching the language type includes:

[0215] For each of the multiple language types, a training sample set for that language type is obtained; the training sample set includes a first type of sample speech using that language type collected from the network.

[0216] Based on the first type of sample speech, the training sample set is expanded to obtain the second type of sample speech using the language type.

[0217] Based on the first type of sample speech and the second type of sample speech, a speech recognition model matching the language type is trained.

[0218] In one possible implementation, the acquisition module is configured as follows:

[0219] The question types of the historical voice data input by the tour guide and the tour guide's evaluation of the historically recommended tour information are obtained to obtain the tour guide's interest data.

[0220] The popularity, ratings, and attributes of attractions in the area to be visited are obtained to acquire attraction data for the area.

[0221] In one possible implementation, the first generation module is configured as follows:

[0222] After preprocessing the data related to the tour object, feature encoding is performed on the preprocessed data of the tour object to obtain the interest features of the tour object;

[0223] After preprocessing the data related to the area to be visited, feature encoding is performed on the preprocessed data of the area to be visited to obtain the scenic spot features of the area to be visited.

[0224] The interest characteristics of the guided object and the scenic spot characteristics of the area to be visited are fused together;

[0225] The fused features are input into the recommendation model to obtain the initial recommended tour information for the tour object; wherein, the initial recommended tour information includes multiple scenic spot tour routes arranged in sequence;

[0226] Obtain the gender, age, occupation type, current physical health status, and emotional state of the guided tour participants;

[0227] Based on at least one of the physiological health status, emotional state, gender attribute, age attribute, and occupation type, the ranking results of the multiple scenic spot tour routes are readjusted to obtain the final recommended tour information for the tour target.

[0228] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0229] It should be noted that the intelligent tour guide device provided in the above embodiments is only illustrated by the division of the above functional modules when performing intelligent tour guidance. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the intelligent tour guide device and the intelligent tour guide method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0230] Figure 5 This is a schematic diagram of the structure of a computer device 500 provided in an embodiment of this application.

[0231] Typically, computer device 500 includes a processor 501 and a memory 502.

[0232] Processor 501 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 501 may be implemented using at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). Processor 501 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In one possible implementation, processor 501 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In another possible implementation, processor 501 may also include an AI (Artificial Intelligence) processor, which handles computational operations related to machine learning.

[0233] Memory 502 may include one or more computer-readable storage media, which may be non-transitory. Memory 502 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In one possible implementation, the non-transitory computer-readable storage media in memory 502 is used to store at least one program code for execution by processor 501 to implement the intelligent navigation method provided in the method embodiments of this application.

[0234] In one possible implementation, the computer device 500 further includes a peripheral device interface 503 and at least one peripheral device. The processor 501, memory 502, and peripheral device interface 503 are connected via a bus or signal line. Each peripheral device is connected to the peripheral device interface 503 via a bus, signal line, or circuit board. The peripheral device includes at least one of the following: radio frequency circuitry 504, display screen 505, camera assembly 506, audio circuitry 507, positioning assembly 508, and power supply 509.

[0235] The peripheral device interface 503 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 501 and the memory 502. In one possible implementation, the processor 501, memory 502, and peripheral device interface 503 are integrated on the same chip or circuit board; in another possible implementation, any one or two of the processor 501, memory 502, and peripheral device interface 503 can be implemented on separate chips or circuit boards, and this application does not limit this.

[0236] The radio frequency (RF) circuit 504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 504 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 504 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. In one possible implementation, the RF circuit 504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 504 can communicate with other terminals via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In one possible implementation, the RF circuit 504 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0237] Display screen 505 is used to display a UI (User Interface). This UI can include graphics, text, icons, videos, and any combination thereof. When display screen 505 is a touch display, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 501 for processing. In this case, display screen 505 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In one possible implementation, there can be one display screen 505, located on the front panel of computer device 500; in another possible implementation, there can be at least two display screens, respectively located on different surfaces of computer device 500 or in a folded design; in yet another possible implementation, display screen 505 can be a flexible display screen, located on a curved or folded surface of computer device 500. Furthermore, display screen 505 can be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 505 can be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0238] The camera assembly 506 is used to acquire images or videos. In one possible implementation, the camera assembly 506 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In one possible implementation, there are at least two rear-facing cameras, which can be any one of a main camera, a depth-sensing camera, a wide-angle camera, or a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In another possible implementation, the camera assembly 506 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash is a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0239] The audio circuit 507 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 501 for processing, or to the radio frequency circuit 504 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, positioned at different locations within the computer device 500. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 501 or the radio frequency circuit 504 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In one possible implementation, the audio circuit 507 may also include a headphone jack.

[0240] The positioning component 508 is used to locate the current geographical location of the computer device 500 in order to enable navigation or LBS (Location Based Service). The positioning component 508 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, Russia's Grenas system, or the European Union's Galileo system.

[0241] Power supply 509 is used to supply power to the various components in computer device 500. Power supply 509 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 509 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0242] Those skilled in the art will understand that Figure 5 The structure shown does not constitute a limitation on the computer device 500, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0243] Figure 6 This is a schematic diagram of the structure of a computer device 600 provided in an embodiment of this application.

[0244] The computer 600 can be a server. The computer device 600 can vary significantly due to differences in configuration or performance, and may include one or more Central Processing Units (CPUs) 601 and one or more memories 602. The memories 602 store at least one line of program code, which is loaded and executed by the processor 601 to implement the intelligent navigation method provided in the various method embodiments described above. Of course, the computer device 600 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The computer device 600 may also include other components for implementing device functions, which will not be elaborated upon here.

[0245] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including program code that can be executed by a processor in a computer device to perform the intelligent navigation method in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0246] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform the above-described intelligent navigation method.

[0247] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0248] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An intelligent navigation method, characterized in that, The method includes: The question types from the historical voice data input by the tour guide and the tour guide's evaluation of the historically recommended tour information are obtained to obtain the tour guide's interest data. Obtain the popularity, ratings, and attributes of attractions in the area to be visited to obtain attraction data for the area to be visited. After preprocessing the interest data, feature encoding is performed on the preprocessed interest data based on the interest analysis model to obtain the interest features of the guided object; after preprocessing the attraction data, feature encoding is performed on the preprocessed attraction data based on the attraction analysis model to obtain the attraction features of the area to be visited. The interest features and the attraction features are fused together; the fused features are input into the recommendation model to obtain the initial recommended tour information for the tour object; wherein, the initial recommended tour information includes multiple attraction tour routes arranged in sequence; The system obtains the gender, age, occupation, current physical health, and emotional state of the tour guide; based on the physical health, emotional state, gender, age, and occupation, it readjusts the sorting results of the multiple scenic spot tour routes to obtain the final recommended tour information for the tour guide. During the tour of the guided object based on the final recommended tour information, in response to obtaining the voice data of the guided object, the language type of the voice data is determined; and based on a voice recognition model that matches the language type, the voice data is used for voice recognition. The response feedback is generated and recognized based on a large language model, and a first voice broadcast type is determined based on the gender attribute and the age attribute; a second voice broadcast type is determined based on the physiological health status and the emotional state; the voice broadcast type that matches the interest data of the guided object in the first voice broadcast type and the second voice broadcast type is determined as the voice broadcast type of the response feedback. The response feedback is broadcast using a language variant that matches the voice broadcast type under the language type; wherein, the language variant includes different language expressions of the language type; The interest analysis model, the scenic spot analysis model, and the recommendation model are large language models. The training process of the interest analysis model, the scenic spot analysis model, and the recommendation model includes: obtaining a training sample set, which includes the interest data of the sample object, the scenic spot data of the sample's visited area, and the historical visit data of the sample object in the sample's visited area; and training the interest analysis model, the scenic spot analysis model, and the recommendation model based on the training sample set.

2. The method according to claim 1, characterized in that, Determining the language type used in the voice data includes: Acoustic feature extraction is performed on the speech data to obtain the filter bank features and speaker pitch features of the speech data; The discrete cosine transform of the filter bank features is used to obtain the Mel-frequency cepstral coefficients of the speech data. The speaker's pitch features and the Mel cepstral coefficients are fused to obtain the fused acoustic features of the speech data. Based on the fusion acoustic features and language type recognition model, the speech data is used to identify the language type and obtain the language type adopted by the speech data.

3. The method according to claim 2, characterized in that, The step of performing language type identification on the speech data based on the fused acoustic features and language type recognition model to obtain the language type adopted by the speech data includes: The fused acoustic features are input into a target neural network; wherein the target neural network is a natural language processing model including an encoder and a decoder; The output features of the encoder are input into the language type recognition model; Based on the output of the language type recognition model, the language type of the speech data is determined.

4. The method according to claim 1, characterized in that, The training process of the speech recognition model matching the language type includes: For each of the multiple language types, a training sample set for that language type is obtained; the training sample set includes a first type of sample speech using that language type collected from the network. Based on the first type of sample speech, the training sample set is expanded to obtain the second type of sample speech using the language type. Based on the first type of sample speech and the second type of sample speech, a speech recognition model matching the language type is trained.

5. An intelligent tour guide device, characterized in that, The device includes: The acquisition module is configured to acquire the question types of the historical voice data input by the tour guide and the tour guide's evaluation of the historically recommended tour information, thereby obtaining the tour guide's interest data. The acquisition module is also configured to acquire the popularity, rating and attributes of attractions in the area to be visited, and obtain the attraction data of the area to be visited. The first generation module is configured to: after preprocessing the interest data, encode the preprocessed interest data based on an interest analysis model to obtain the interest features of the tour object; after preprocessing the attraction data, encode the preprocessed attraction data based on an attraction analysis model to obtain the attraction features of the area to be visited; fuse the interest features and the attraction features; input the fused features into a recommendation model to obtain the initial recommended tour information for the tour object; wherein the initial recommended tour information includes multiple attraction tour routes arranged in order; acquire the gender attribute, age attribute, occupation type, current physical health status, and emotional state of the tour object; and readjust the sorting results of the multiple attraction tour routes according to the physical health status, emotional state, gender attribute, age attribute, and occupation type to obtain the final recommended tour information for the tour object. The determination module is configured to, in response to acquiring the voice data of the guided object during the tour based on the final recommended tour information, determine the language type of the voice data. The speech recognition module is configured to perform speech recognition on the speech data based on a speech recognition model that matches the language type. The second generation module is configured to generate and recognize response feedback based on a large language model and text matching. The output module is configured to: determine a first voice broadcast type based on the gender attribute and the age attribute; determine a second voice broadcast type based on the physiological health state and the emotional state; determine the voice broadcast type that matches the interest data of the guided object between the first voice broadcast type and the second voice broadcast type as the voice broadcast type for the response feedback; and broadcast the response feedback using a language variant that matches the voice broadcast type under the language type; wherein the language variant includes different language expressions of the language type. The interest analysis model, the scenic spot analysis model, and the recommendation model are large language models. The training process of the interest analysis model, the scenic spot analysis model, and the recommendation model includes: obtaining a training sample set, which includes the interest data of the sample object, the scenic spot data of the sample's visited area, and the historical visit data of the sample object in the sample's visited area; and training the interest analysis model, the scenic spot analysis model, and the recommendation model based on the training sample set.

6. A computer device, characterized in that, The device includes a processor and a memory, the memory storing at least one line of program code, which is loaded and executed by the processor to implement the intelligent navigation method as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the intelligent navigation method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Personalized tour route recommendation method

    CN108829852A

  • Intelligent voice guide method, device and equipment and storage medium

    CN110459203A

  • Voice interaction method and system, terminal equipment and storage medium

    CN115019788A

  • Voice processing method and device, equipment and storage medium

    CN116312477A