Automatic user language detection for content selection
By combining account profiles and browsing history, and using natural language processing and machine learning models, the system dynamically selects content items familiar to multilingual users, solving the problem of inaccurate language recognition in existing technologies and achieving more efficient content selection and a better user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-14
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies struggle to accurately identify the language preferences of multilingual users when automatically detecting their language, leading to inaccurate content selection, increased consumption of computing and network resources, and reduced quality of human-computer interaction.
By combining the client device's account profile, browsing history, and search query keywords, and using natural language processing and machine learning models, the system predicts the multiple languages the user uses and dynamically selects content items that match the user's familiar languages.
It improves the accuracy of user language prediction, reduces computational and network resource consumption, enhances the quality of human-computer interaction, and provides a larger pool of comprehensible content items.
Smart Images

Figure CN115176242B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of data processing, and, more specifically, to systems and methods for automatically detecting user language for content selection. BACKGROUND
[0002] In a computer networking environment such as the Internet, content providers can provide content items to be inserted into information resources (e.g., web pages) that are processed and rendered by applications (e.g., web browsers) running on client devices. SUMMARY
[0003] At least one aspect relates to systems and methods for automatically detecting user language for content selection. A data processing system having one or more processors coupled with memory can receive, from a client device, a request for content that identifies an account profile. The data processing system can determine, from a plurality of languages, a first set of candidate languages using log records of a plurality of activities of the account profile. The data processing system can identify, responsive to the request for content, a plurality of information resources to be presented to the client device according to rankings. The data processing system can determine, from the plurality of languages, a second set of candidate languages based on content in each information resource and a corresponding ranking of each information resource. The data processing system can identify a set of languages included in both the first set of candidate languages and the second set of candidate languages, the set of languages including a first language and a second language from the plurality of languages. The data processing system can store, in one or more data structures, an association between the account profile, the first language, and the second language.
[0004] In some implementations, the data processing system can generate a first confidence score for the first language based on a first number of occurrences of the first language in at least one of the plurality of information resources or the plurality of activities identified responsive to the request for content. In some implementations, the data processing system can generate a second confidence score for the second language based on a second number of occurrences of the second language in at least one of the plurality of information resources or the plurality of activities identified responsive to the request for content.
[0005] In some implementations, the data processing system can include the first language in the first set of candidate languages responsive to determining that the first confidence score for the first language is greater than a threshold score. In some implementations, the data processing system can include the second language in the second set of candidate languages responsive to determining that the second confidence score for the second language is greater than the threshold score.
[0006] In some implementations, the data processing system can identify a first plurality of content items in a first language and a second plurality of content items in a second language. In some implementations, the data processing system can provide, to the client device, a content item selected from one of the first plurality of content items and the second plurality of content items, the content item being in one of the first language or the second language.
[0007] In some implementations, the data processing system can identify a selection value for each content item of the first plurality of content items in the first language and the second plurality of content items in the second language. In some implementations, the data processing system can select, from the first plurality of content items and the second plurality of content items, a content item to provide to the client device according to a content selection protocol, the content item being in one of the first language or the second language.
[0008] In some implementations, the data processing system can identify an information resource associated with the content item in the first language or the second language. In some implementations, the data processing system can determine that the language of the content item corresponds to the language of the information resource. In some implementations, the data processing system can add the content item to the plurality of candidate content items for selection by the client device in response to determining that the language of the content item corresponds to the language of the information resource.
[0009] In some implementations, the data processing system can identify a third set of candidate languages from at least one of a language setting of the account profile, a language configuration of an application running on the client device, or one or more keywords included in the request for content. In some implementations, the data processing system can identify a set of languages included in the first set of candidate languages, the second set of candidate languages, and the third set of candidate languages.
[0010] In some implementations, the data processing system can determine, from the plurality of languages, a plurality of activities including at least one of a search query received from the client device, an access to an information resource by the client device, and an interaction with an element on the information resource based on the plurality of activities identified in the log records. In some implementations, the data processing system can determine, from the plurality of languages, based on a frequency of each language of the second set of candidate languages identified across a plurality of information resources in response to the request for content.
[0011] In some implementations, the data processing system can receive a query including one or more keywords from a client device. In some implementations, the data processing system can perform a search operation using the one or more keywords of the query to identify a plurality of information resources. In some implementations, the data processing system can provide an output including at least one of the plurality of information resources and a content item selected from one of a first plurality of content items in a first language and a second plurality of content items in a second language, the content item being in one of the first language or the second language.
[0012] These and other aspects and implementations are discussed in detail below. The foregoing information and the following detailed description include illustrative examples of various aspects and implementations, and provide an overview or framework for understanding the nature and character of the claimed aspects and implementations. The drawings provide one explanation of aspects and implementations and are not intended to be a detailed BRIEF DESCRIPTION OF DRAWINGS
[0013] The drawings are not intended to be drawn to scale. Like references and designations in various drawings indicate like elements. For purposes of clarity, not every component can be called out in every drawing. In the drawings:
[0014] Figure 1 is a block diagram of a system for automatically detecting user language for content selection in accordance with an illustrative embodiment;
[0015] Figure 2 is a sequence diagram of a query processing process of a system for automatically detecting user language for content selection in accordance with an illustrative embodiment;
[0016] Figure 3 is a sequence diagram of a language profiling process of a system for automatically detecting user language for content selection in accordance with an illustrative embodiment;
[0017] Figure 4 is a sequence diagram of a result evaluation process of a system for automatically detecting user language for content selection in accordance with an illustrative embodiment;
[0018] Figure 5 is a sequence diagram of a content selection process of a system for automatically detecting user language for content selection in accordance with an illustrative embodiment;
[0019] Figure 6 is a sequence diagram of a result provision process of a system for automatically detecting user language for content selection in accordance with an illustrative embodiment;
[0020] Figure 7is a flowchart of a method of automatically detecting a user language for content selection according to an illustrative embodiment; and
[0021] Figure 8 is a block diagram illustrating a general architecture for a computer system, which can be employed to implement elements of the systems and methods described and illustrated herein, according to an illustrative implementation. DETAILED DESCRIPTION
[0022] The following is a more detailed description of various concepts and implementations related to methods, apparatus, and systems for determining a user language in a networked environment. The various concepts introduced above and discussed in greater detail below can be implemented in any of numerous ways, as the described concepts are not limited to any particular implementation.
[0023] A centralized service of a content distribution platform can select content items from various content providers for transmission to a client device using any number of selection parameters. Each content item can have been configured to present audio, video, or textual content in one particular language (e.g., English). The selection parameters for each content item can be set by the respective content provider to define when the content item is to be provided to a client device when associated with a particular language identifier. When a request or query for content is received from a client device, the service can identify a language used by a user of the client device. The language can be identified from a language setting of an account associated with the user, a language configuration of an application (e.g., web browser) on the client device, or text of the query itself. With this identification, the service can select and provide one of the content items having content in the same language as the language identified for the client device in response to the request. For example, when the language identified as used by the user of the requesting client device is also Italian, the service can provide a content item having video content in Italian as specified in the selection parameters.
[0024] One drawback of selecting content items in this manner can be that this approach ignores the possibility that the user of the requesting client device can be multilingual (e.g., Spanish and Italian). This oversight can be further exacerbated by the fact that many users, including a vast majority of multilingual users, do not self-report which languages they use in account profiles or application settings. Another drawback of this approach can be that the accuracy of identifying other languages used by the user is very low, even if the received query is in a different language. This can be because the text of the query is typically very short and therefore ambiguous in giving limited context, where a keyword in the query can be a word used in multiple languages. For example, a query containing the keyword "taxi" can be ambiguous because it is difficult to determine whether the user wants the language to be English or French, or another language, as all of these languages use this word.
[0025] As a result, for multilingual users (e.g., both Spanish and Italian), the set of candidate content items for potential selection may be limited to one language (e.g., Spanish or Italian), thus excluding other languages the user may be familiar with or proficient in. Excluding such content items in other languages can lead to greater consumption of computational and network resources, as users may repeatedly query to find relevant content. Furthermore, excluding content items from other languages may also result in lower quality human-computer interaction (HCI) between the user and the client device, since the content may only be in one language, but not in other languages the user is familiar with.
[0026] To address these and other technical challenges, content delivery platform services can determine the language used by a user on a requesting client device based on a mixture of various signals of varying quality and coverage. The service can identify the user's declared language from account or application settings and can also derive the language from keywords in the query itself. In addition to these factors, the service can construct a user language profile from the client device's browsing history. The service can identify various access activities performed via the client device, such as those identified in the browsing history. These activities can include, for example, accessing information resources (e.g., web pages), entering input (e.g., comments) on the graphical user interface of an information resource, and previous queries leading to that resource. Through this identification, the service can determine the language associated with the access activities to construct the user language profile. The service can also consider the language identified from keywords in the user's declared language and the query itself within the user language profile. The user language profile can indicate that the user of the client device is predicted to use one or more languages.
[0027] Combined, the service can identify the language used by the user on the client device from the search results of a query. In this identification, the service can perform a web search operation using one or more keywords from the query to find a set of information resources whose content matches or is related to the keywords. The web search operation may involve using or invoking a search engine with the query and returning the set of information resources as search results. The set of information resources can be sorted sequentially based on a ranking indicating the relevance of the resulting information resources to the query keywords. The service can parse each information resource to determine the language from its content. The service can narrow down the number of languages by considering the ranking of information resources from which the derived language is derived and the frequency of the determined language among the information resources in the search results.
[0028] The service can identify a set of candidate content items for each identified language from an initial set of languages indicated in the constructed user language profile. Each content item can have a selection parameter indicating that the content item is to be selected when the language determined for the user matches a language defined by the content provider. The service can filter the set of languages in the user profile by identifying an intersection between the set of languages in the user profile and the set of languages determined from the search results. By filtering the number of languages predicted for the user, the service can expand the set of candidate content items that are eligible to be selected for providing to the client device.
[0029] Once the content items are filtered, the service can run a content selection process to select content items to provide to the client device. This can result in selecting content items in a language that is different from the language declared by the user in the account profile or application settings. For example, a client device submitting a query can have an account profile set to indicate that the user uses English, but the browsing history can indicate that the user frequently visits Polish web pages. From the access history and the search results, it can be determined that the user understands both English and Polish, and content items using either language can be selected for the pool of eligible content items. From the content selection process, the service can also select content items using either language. The content items provided to the client device can be presented with the search results found using the keywords of the query. The language of the provided content items can use a different language than at least some of the search results and the same language as some other search results.
[0030] By using multiple factors in this way, the accuracy of predicting the language used by the user can be significantly increased, up to 70-90%, compared to using only the language declared by the user or derived from the keywords of the query. In addition, the set of content items to select and provide from can be expanded to include multiple languages determined with higher accuracy and precision. Including these content items for selection can result in reduced consumption of computing and network resources, and fewer queries by the user through the client device to find relevant content. In combination with the increased accuracy of predicting the language, adding content items across multiple languages can result in a higher quality HCI between the user and the client device, as content can be provided in any language determined to be known by the user.
[0031] Because the first language and the second language have been determined based on both the account profile's activity and information resources identified in response to requests for content, the selected content item is in a language familiar to the user. Thus, by virtue of the methods disclosed herein, a greater pool of intelligible content items is provided with high accuracy. The selection of content items provided to the client device is also improved because it is drawn from content items in more than one language. Thus, the method effectively selects content items that the user can understand because the user's proficiency in more than one language is accurately determined. As a result, the selection of intelligible content data can be made from a greater pool of data than provided by existing methods.
[0032] Reference will now be made to Figure 1 depicted is a block diagram depicting one embodiment of a computer networking environment or system 100 for determining a user's language. Generally, the system 100 can include at least one network 105 for communication between components of the system 100. The system 100 can include at least one data processing system 110 to process requests communicated via the network 105. The data processing system 110 can include at least one query handler 135, at least one profile extractor 140, at least one search evaluator 145, at least one language evaluator 150, and at least one content aggregator 155, among others. The system 100 can include at least one content provider 115 to provide content items. The system 100 can include at least one content publisher 120 to provide information resources (e.g., web pages). The system 100 can include at least one client device 125 to communicate via the network 105. The system 100 can include at least one index service 130 (sometimes referred to herein as a search engine and web crawler) to find information resources using one or more keywords provided by the client device 125. Each component of the system 100 (e.g., the network 105, the data processing system 110 and its components, the content provider 115 and its components, the content publisher 120 and its components, and the client device 125 and its components) can be implemented using the components of the computing system 800 described in detail herein in connection with Figure 8 The components of the computing system 800 described in detail herein can be implemented as physically separate elements or can be implemented as a single element, such as a system on a chip (SoC). The components of the computing system 800 described in detail herein can be implemented using one or more integrated circuits (ICs) or other hardware components.
[0033] In more detail, the network 105 of the system 100 can communicatively couple the data processing system 110, the content providers 115, the content publishers 120, and the client devices 125 to each other. Each of the data processing system 110, the content providers 115, and the content publishers 120 of the system 100 can include a plurality of servers located in at least one data center or server farm that are communicatively coupled to each other via the network 105. The data processing system 110 can communicate with the content providers 115, the content publishers 120, and the client devices 125 via the network 105. The content providers 115 can communicate with the data processing system 110, the content publishers 120, and the client devices 125 via the network 105. The content publishers 120 can communicate with the data processing system 110, the content publishers 120, and the client devices 125 via the network 105. The client devices 125 can communicate with the data processing system 110, the content providers 115, and the content publishers 120 via the network 105.
[0034] The content providers 115 can include servers or other computing devices operated by content provider entities to provide content items for display on information resources at the client devices 125. The content provided by the content providers 115 can take any convenient form. For example, third-party content can include content related to other displayed content and can be, for example, a website page related to the displayed content. The content can include a third-party content item or creative (e.g., an advertisement) for display on an information resource, such as an information resource that includes primary content provided by the content publishers 120. The content items can also be displayed on search result web pages. For example, the content providers 115 can provide or be the source of content items for display in content slots (e.g., iframe elements) of information resources, such as a company's web page where the primary content of the web page is provided by the company, or for display on a search engine-provided search results login page. The content items associated with the content providers 115 can be displayed on information resources other than web pages, such as content displayed as part of the running of an application on a smartphone or other client device 125.
[0035] Content publisher 120 can include a server or other computing device operated by a content publishing entity to provide information resources including primary content for display via network 105. For example, content publisher 120 can include a web page operator that provides primary content for display on information resources. An information resource can include content other than content provided by content publisher 120, and an information resource can include a content slot configured for display of a content item from content provider 115. For example, content publisher 120 can operate a website for a company, and can provide content about the company for display on web pages of the website. A web page can include a content slot configured for display of a content item provided by content provider 115 or provided by content publisher 120 itself. In some implementations, content publisher 120 can include a search engine computing device (e.g., a server) operated by a search engine operator that operates a search engine website. Primary content of a search engine web page (e.g., a results or landing web page) can include results of a search and third-party content items, such as content items from content provider 115, displayed in content slots of the information resource.
[0036] Data processing system 110 can include a server or other computing device operated by a content placement entity to select or identify content items to be inserted into content slots of information resources via network 105. In some implementations, data processing system 110 can include servers and computing devices operated by a search engine operator. In some implementations, data processing system 110 can include a content placement system (e.g., an online advertising server or other data provider). Data processing system 110 can maintain a list of content items from which to select for insertion into content slots of information resources provided over network 105. The list can be maintained on a database accessible to data processing system 110. Content items or identifiers (e.g., addresses) of content items can be provided by content provider 115. In some implementations, data processing system 110 can include search engine computing devices (e.g., servers) operated by a search engine operator that operates a search engine website. Primary content of a search engine web page (e.g., a results or landing web page) can include results of a search and third-party content items, such as content items from content provider 115, displayed in content slots of the information resource.
[0037] Each client device 125 can include a computing device to communicate over the network 105 to display data. The display content can include content (e.g., information resources) provided by the content publishers 120 and content (e.g., content items for display on content slots of information resources) provided by the content providers 115, as identified by the data processing system 110. The client devices 125 can include desktop computers, laptop computers, tablet computers, smartphones, personal digital assistants, mobile devices, consumer computing devices, servers, clients, digital video recorders, television set-top boxes, video game consoles, or any other computing devices configured to communicate via the network 105.
[0038] The index service 130 can include servers or other computing devices operated by a search engine service to aggregate information resources accessible via the network 105 and provide search results in response to queries of the client devices 125. In some implementations, the index service 130 can be part of the data processing system 110 or the content publishers 120. In some implementations, the functionality of the index service 130 can be distributed across one or more of the data processing system 110, the content publishers 120, or the index service 130. The primary content of a search engine webpage (e.g., a results or landing webpage) can include search results as well as third-party content items displayed in content slots of information resources, such as content items from the content providers 115.
[0039] The client devices 125 can be operated or used (e.g., using input / output (I / O) devices) by at least one user 160. In some implementations, the user 160 can be associated with the client device 125A (e.g., logged into the client device 125A via an account). The user 160 can be proficient or can understand multiple languages, such as a first language 165A and a second language 165B (hereinafter collectively referred to as languages 165). The languages 165 can include any natural language, such as English, Spanish, French, German, Mandarin, Urdu, Arabic, Russian, Portuguese, Japanese, Korean, Indonesian, and Italian, among others. The languages 165 can be represented in text (e.g., using symbols). The user 160 can also be proficient or understand one language, such as the first language 165A or the second language 165B.
[0040] Reference is now made to Figure 2FIG. 2B depicts a sequence diagram of a query processing process 250 for automatically detecting a user language for content selection, according to at least one embodiment. As shown, the client device 125 can run or include at least one application 205. The application 205 can be a program executable on the client device 125 to access resources via the network 105. For example, the application 205 can be a web browser, a web application, a mobile application, or a word processing application, among others. The application 205 can have retrieved or fetched at least one information resource 210 (e.g., a webpage) from the data processing system 110 or the content publisher 120. The information resource 210 can include one or more user interface elements with which the user 160 can interact via I / O devices of the client device 125 to input. In some embodiments, the information resource 210 can correspond to a search engine webpage from the data processing system 110. The search engine webpage can include at least one user interface element (e.g., a text box) to input a query for searching content. The input to the user interface element of the information resource 210 can be in a first language or a second language.
[0041] The application 205 can have or be associated with at least one language configuration 215 (sometimes referred to herein as a language setting). The language configuration 215 can define, specify, or otherwise identify one or more languages to be used on the application 205. In accordance with the language configuration 215, the application 205 can send a request for content in a specified language and retrieve one or more information resources in the specified language (e.g., the information resource 210). For example, the language configuration 215 can specify that Portuguese is to be used. In this example, the application 205 can retrieve a Portuguese webpage by sending a request for content in Portuguese. In some embodiments, the language configuration 215 of the application 205 can be set to a default language. The default language can be based on a geographic region of the client device 125, a language setting of the client device 125 (e.g., specified by an operating system (OS)), or preconfigured by the application 205.
[0042] Additionally, the application 205, the client device 125, or the user 160 can be associated with at least one account profile 220. The account profile 220 can correspond to or be associated with an account with which the user 160 is authenticated to use the client device 125 or the application 205. For example, the user 160 can log in to use the application 205 using an account identifier and a password for the account to be logged in. The account profile 220 can be associated with the user 160 via the account identifier. The account profile 220 can be maintained on the client device 125 or a remote service (e.g., the data processing system 110) accessible via the application 205. The account profile 220 can define, specify, or otherwise identify one or more languages associated with the user 160 (or an extended account), the client device 125, or the application 205 (e.g., via a language setting of the account). As with the language configuration 215, the language specified by the account profile 220 can be used to send requests for content and retrieve one or more information resources (e.g., the information resource 210).
[0043] The application 205 running on the client device 125 can generate and transmit at least one request 225 for content to the data processing system 110 over the network 105. The generation and transmission of the request 225 can be in response to input by the user 160 via the application 205 (e.g., a user element) running on the client device 125. The request 225 can identify the account profile 220. In some implementations, the request 225 can include a reference to an identifier of the user 160 (e.g., a set of alphanumeric characters in a specified field), the associated account, or the account profile 220. In some implementations, the request 225 can include or can correspond to a search query generated via a search engine webpage. For example, the request 225 can be generated upon input of a query on a search engine webpage loaded on the application 205. In some implementations, the request 225 can include or identify the language configuration 215 associated with the application 205 or the client device 125. For example, the request 225 can include one or more languages indicated by the language configuration 215.
[0044] The request 225 can include one or more keywords 230A-N (hereinafter collectively referred to as keywords 230). Input of the one or more keywords 230 of the request 225 can be performed via one or more I / O devices of the client device 125. The one or more keywords 230 of the query can correspond to or include a set of alphanumeric characters in a text input. In some implementations, the keywords 230 of the query can correspond to input on an element of the information resource 210 (e.g., a search engine). In some implementations, the input can be audio input via a microphone or another form of transducer for audio input. The one or more keywords 230 of the query can correspond to a portion of the audio input that corresponds to the set of alphanumeric characters. In some implementations, the application 205 can use natural language processing (NLP) techniques (e.g., speech recognition) to convert the input audio to a set of alphanumeric characters (e.g., text) to be included as keywords 230 of the query. In some implementations, the input audio can be included in the query to be converted to a set of alphanumeric characters at the data processing system 110.
[0045] The query processor 135 running on the data processing system 110 can retrieve, identify, or otherwise receive the request 225 from the client device 125. Upon receipt, the query processor 135 can parse the request 225 to identify the keywords 230. In some implementations, the query processor 135 can extract a text input included or identified in the request 225. Using the extracted text, the query processor 135 can determine or identify the one or more keywords 230. For example, the query processor 135 can group or identify sets of alphanumeric characters separated from one another by spaces or line breaks as keywords 230 of the request 225. In some implementations, the query processor 135 can extract an audio input included or identified in the request 225. The query processor 135 can apply NLP techniques (e.g., speech recognition) to identify keywords 230 from one or more portions of the audio input of the request 225. In applying the NLP techniques, the query processor 135 can build, train, and maintain a speech recognition model to apply to audio to identify keywords 230.
[0046] Using information associated with or identified by the request 225, the query processor 135 running on the data processing system 110 can determine or identify candidate languages 235A-N (hereinafter collectively referred to as candidate languages 235) of the candidate set 240. The candidate languages 235 can be one or more languages that the user 160 is estimated, predicted, or otherwise determined to use. The information associated with the request 225 can include the language configuration 215, the account profile 220, and the keywords 230. In some implementations, the query processor 135 can determine or identify the candidate languages 235 based on the language configuration 215 associated with the application 205 or the client device 125. The query processor 135 can parse the request 225 to identify one or more languages defined by the language configuration 215 as candidate languages 235. The query processor 135 can add, insert, or include the candidate languages 235 identified from the language configuration 215 to the candidate set 240.
[0047] In some implementations, the query processor 135 can determine or identify the set of candidate languages 235 based on the account profile 220. The query processor 135 can parse the request 225 to identify the account profile 220. For example, the query processor 135 can parse the request 225 to extract an account identifier associated with the account profile 220 and can find the account profile 220 associated with the account identifier. From the account profile 220, the query processor 135 can identify one or more languages defined as used by the user 160. The query processor 135 can add, insert, or include the candidate languages 235 identified from the account profile 220 to the candidate set 240.
[0048] In some implementations, at least one language identification model 245 can be established and maintained by the data processing system 110 to determine a language used in the keywords 230 of the request 225. The language identification model 245 can be an artificial intelligence (AI) algorithm or a machine learning (ML) model (e.g., an artificial neural network, an n-gram model, a Bayesian network, a random forest, a support vector machine, or a decision tree, etc.). In general, the language identification model 245 can include a set of inputs, a set of outputs, and a set of weights (sometimes referred to as parameters herein) to relate the inputs and the outputs. The inputs can include text (e.g., the keywords 230 extracted from the request 225). The outputs can include or identify the language 235 in which the text is used. In some implementations, the outputs can also include a likelihood measure indicating a confidence of the text in each language 235. The weights can be according to the architecture of the AI algorithm or the ML model.
[0049] The language recognition model 245 can be trained using a training dataset (e.g., by the data processing system 110). The training can be according to a supervised or unsupervised learning algorithm. The training dataset can include a corpus of text for each language 235 that is labeled for the corpus. By applying the text from each corpus to the language recognition model 245, a result corresponding to one of the languages 235 can be generated from the language recognition model 245. Based on a comparison of the result to the labeled language of the corpus of the dataset in the training, an error can be determined. The error can be a mean squared error (MSE), a root mean squared error (RMSE), or a cross-entropy error, among others. Using the error, the weights of the language recognition model 245 can be adjusted or modified. The updating of the weights of the language recognition model 245 can be repeated until convergence. For example, the weights of the language recognition model 245 can be determined to have converged when a change in the weight values is determined to be less than a convergence threshold. The establishment and training of the language recognition model 245 can be performed prior to receiving the request 225 from the one or more client devices 125.
[0050] In some implementations, the query processor 135 can identify or determine the candidate language 235 based on one or more of the keywords 230 of the request 225. The first language can refer to the language used in the keywords 230 of the request 225. To determine, in some implementations, the query processor 135 can apply the language recognition model 245 to the keywords 230 of the request 225. In the application, the query processor 135 can feed the keywords 230 of the request 225 as input to the language recognition model 245. The query processor 135 can use the weights of the language recognition model 245 to process the input to generate or produce an output. The output of the language recognition model 245 can indicate which language 235 is used in the keywords 230 of the request 225. In some implementations, the output can include the languages 235 with corresponding likelihood metrics. The query processor 135 can identify the language 235 from the output generated by the language recognition model 245. In some implementations, the query processor 135 can identify the language 235 with the highest likelihood metric computed by the language recognition model 245. The query processor 135 can add or insert the candidate language 235 determined using the language recognition model 245 to the candidate set 240.
[0051] Reference is now made to Figure 3FIG. 10 depicts a sequence diagram of a language profiling process of the system 100 for automatically detecting a user language for content selection. As shown, from at least one database 330, the profile extractor 140 can select or identify at least one log record 305 for an account profile 220 identified by a request 225. The log record 305 can be maintained and stored on the database 300. The log record 305 includes or identifies one or more activities 310A-N (hereinafter collectively referred to as activities 310). In some embodiments, the activities 310 of the log record 305 can be arranged using one or more data structures. For example, the log record 305 can be maintained using a relational database maintained with a database management system (DBMS) and can include an entry for each activity 310 of the log record 305.
[0052] The log record 305 can be maintained on the database 300 for a particular client device 125, a particular application 205, or a particular account profile 220 (e.g., as depicted). The activities 310 identified in the log record 305 can correspond to previous actions performed by the client device 125 (or application 205) associated with the account profile 220 via the network 105. The activities 310 can also be associated with or include content. In some embodiments, at least one activity 310 of the log record 305 can include or correspond to a request (e.g., a search query) for content received from the client device 125. For example, a search query including a keyword can have been submitted from the client device 125 associated with the account profile 220 to retrieve a web page using the keyword. In some embodiments, at least one activity 310 of the log record 305 can include or correspond to an access of an information resource (e.g., a web page) by the client device 125. For example, a cookie can be used to identify a web page accessed by the client device 125 associated with the account profile and the access of the client device 125 can be recorded on the log record 305. In some embodiments, at least one activity 310 of the log record 305 can include or correspond to an interaction with an element on an information resource performed via the client device 125. For example, a user 165 associated with the account profile 220 can enter a comment on a web page and the comment can be identified by an activity 310 recorded on the log record 305.
[0053] Using one or more activities 310 of the log records 305, the profile exporter 140 can select, identify, or determine one or more candidate languages 235'A-N (hereinafter collectively referred to as candidate languages 235') of the candidate set 240'. In some implementations, the profile exporter 140 can select or identify a subset of activities 310 for use in determining the candidate languages 235' of the candidate set 240'. For example, the profile exporter 140 can select a subset of activities 310 from a time window prior to receiving the request 225. For each activity 310 identified from the log records 305, the profile exporter 140 can identify or determine a candidate language 235'. In determining, the profile exporter 140 can parse the activity 310 to identify an action performed by the client device 125 (or the application 205) via the network.
[0054] With this identification, the profile exporter 140 can identify content associated with the action corresponding to the logged activity 310. The content can include, for example, keywords in a request for content, text on an information resource visited, and input on one or more user interface elements on the information resource, among others. The profile exporter 140 can apply the language identification model 245 to the content associated with the activity to determine the candidate language 235' in the manner described above. The process of identifying activities 310 and determining candidate languages 235' from content associated with the activities 310 can be repeated through the log records 305.
[0055] For each candidate language 235' identified from the activities 310, the profile exporter 140 can compute, determine, or otherwise generate a confidence score. The confidence score can indicate a probability or degree of certainty that the user 165 actually uses the corresponding candidate language 235'. In computing, the profile exporter 140 can identify a number of occurrences of the candidate language 235' from the activities 310 of the log records 305. In some implementations, the profile exporter 140 can maintain a counter to track the number of occurrences of the candidate language 235' identified from parsing the activities 310 of the log records 305. Based on the number of occurrences, the profile exporter 140 can generate the confidence score. In some implementations, the profile exporter 140 can use a frequency of occurrence of the corresponding language 235' to determine the confidence score. The frequency can be based on the number of occurrences of the corresponding candidate language 235' and a total number of occurrences of all identified candidate languages 235'. Generally, the more occurrences, the higher the confidence score can be. Conversely, the fewer occurrences, the lower the confidence score of the corresponding candidate language 235' can be.
[0056] Using the confidence scores, the profile exporter 140 can determine whether to add or include candidate languages 235' in the candidate set 240'. In some implementations, the profile exporter 140 can select candidate languages 235' corresponding to the N highest confidence scores to include to the candidate set 240'. In some implementations, the profile exporter 140 can compare the confidence score of a corresponding candidate language 235' to a threshold score to determine whether to include to the candidate set 240'. The threshold score can delineate or demarcate a value of the confidence score of the corresponding candidate language 235' to be included in the candidate set 240'. When the confidence score satisfies (e.g., is greater than or equal to) the threshold score, the profile exporter 140 can select the corresponding candidate language 235' to include to the candidate set 240'. On the other hand, when the confidence score does not satisfy (e.g., is less than) the threshold score, the profile exporter 140 does not select the corresponding candidate language 235' to include in the candidate set 240'.
[0057] Beneficially, by using the confidence scores to determine whether to include candidate languages to the candidate set, the number of different languages in the candidate set can only increase when the candidate language has been sufficiently used (i.e., satisfies or exceeds the threshold score). For example, when the confidence score does not satisfy or exceed the threshold score, the use of the candidate language has been determined to be insufficient for such language to be part of the candidate set. As a result of using the threshold score in this manner, the candidate set can only include languages that the particular user actually understands. As a result, this avoids increasing the size of the candidate set, and thus, avoids unnecessarily increasing the number of content items that the method can identify. By accurately limiting the number of content items to only those languages that the user is able to use, the method is able to more efficiently provide content items in a language that the user can understand.
[0058] Referring now to Figure 4 , a sequence diagram of a result evaluation process 400 for the system 100 for automatically detecting user language for content selection is depicted. As shown, the search evaluator 145 running on the data processing system 110 can use the keywords 230 of the request 225 to conduct, run, or otherwise perform at least one search operation 405 to identify at least one query result 410. To perform the search operation 405, the search evaluator 145 can invoke the indexing service 130 using the keywords of the keywords 230 of the request 225. In some implementations, the search evaluator 145 can send or provide the keywords 230 by forwarding a request 225' (also referred to herein as a query). The request 225' can include at least a subset of the keywords 230 of the original request 225. In some implementations, the search evaluator 145 can generate and send the request 225' including the keywords 230 of the original request 225 to the indexing service 130.
[0059] The index service 130 can aggregate one or more information resources (e.g., web pages) that are accessible via the network 105 (e.g., the Internet). In some implementations, the index service 130 can conduct or perform an indexing process (also referred to herein as web indexing or spidering) over the network 105 to identify information resources 420A-N (hereinafter collectively referred to as information resources 420). Each information resource 420 can be uniquely identified or referenced by an identifier (e.g., a uniform resource locator (URL)). In addition, each information resource 420 can include content (e.g., textual or audiovisual) and can be associated with metadata. The index service 130 can parse each identified information resource 420 to extract or identify at least a portion of the content included in the information resource 420 and the metadata associated with the information resource 420. With this identification, the index service 130 can maintain and store the identifier of the information resource 420, the at least the portion of the content, and the metadata on the database 410.
[0060] Upon receipt, the index service 130 can parse the request 225' (or the request 225) to extract or identify one or more keywords 230'. Using the keywords 230', the index service 130 can identify one or more information resources 420. In some implementations, the index service 130 can use the keywords 230' to search the database 410 to find one or more information resources 420 that were aggregated via the indexing process. In the identification, the index service 130 can compare the keywords 230' from the request 225' to the content or metadata of the information resources 420. In some implementations, the index service 130 can use or apply a natural language processing (NLP) process to compare the keywords 230' to the content or metadata of the information resources 420. For example, the index service 130 can use a semantic knowledge graph to generate additional words and phrases with semantic similarity (e.g., synonyms) as the keywords 230' of the request 225'. The index service 130 can then use the additional keywords or phrases to match the content or metadata of the information resources 420. Based on the comparison, the index service 130 can determine whether at least a portion of the content or metadata of the information resources 420 matches or corresponds to the one or more keywords 230'. In some implementations, the index service 130 can determine that the information resources 420 include content or metadata that matches the keywords 230' or additional associated words and phrases.
[0061] Based on the determination, the index service 130 can generate at least one query result 415 to provide to the search evaluator 145. The query result 415 can include or identify one or more information resources 420 that were determined to have content or metadata that matches or corresponds to the keywords 230' of the request 225'. When the content or metadata of an information resource 402 is determined not to match or correspond to any of the keywords 230', the index service 130 can exclude the information resource 420 from the query result 415. Conversely, when the content or metadata of an information resource 402 is determined to match or correspond to a keyword 230', the index service 130 can add or include the information resource 420 to the search query 415.
[0062] With the identification of the one or more information resources 420 to include, the index service 130 can determine or generate at least one ranking 425 of the query result 415. The ranking 425 can specify, define, or identify a relevancy of the information resources 420 to the keywords 230' of the request 225'. The ranking 425 can also identify an order in which the information resources 420 (or identifiers of the information resources 420) are to be presented (e.g., on a search results page). In determining, the index service 130 can calculate, determine, or generate a relevancy score for each identified information resource 420. The calculation of the relevancy score can be based on a number of occurrences of the keywords 230' in the content or metadata of the information resource 420. Based on the relevancy scores of the identified information resources 420, the index service 130 can determine the ranking 425. Generally, the higher the relevancy score of a given information resource 420 in the query result 415, the higher the information resource 420 can be in the ranking 425. Conversely, the lower the relevancy score of a given information resource 420 in the query result 415, the lower the information resource 420 can be in the ranking 425. With the generation, the index service 130 can send or provide the query result 415 to the search evaluator 145.
[0063] The search evaluator 145 can identify the information resources 420 ordered according to the rankings 425 from the search operation 405. In some implementations, the search evaluator 145 can parse the query results 415 received from the index service 130 to identify the information resources 420 and the rankings 425. Based on the information resources 420 and the rankings 425, the search evaluator 145 can select, identify, or determine one or more candidate languages 235''A-N (hereinafter collectively referred to as candidate languages 235'') of the candidate set 240''. For each information resource 420, the search evaluator 145 can identify or determine a candidate language 235'' used by the information resource 420. The search evaluator 145 can parse the information resource 420 to extract or identify at least a portion of the content. The search evaluator 145 can apply the language identification model 245 to the content of the information resource 420 to determine the candidate language 235'' in the manner described above. The process of identifying the information resources 420 and the candidate languages 235'' can be repeated through the query results 415.
[0064] In some implementations, the search evaluator 145 can use the candidate set 240' to arrange and generate the candidate set 240''. The search evaluator 145 can use the candidate languages 235' in the candidate set 240' as an initial set of candidate languages 235'' of the candidate set 240''. When the candidate language 235' is determined in one or more information resources 420 of the query results 415, the search evaluator 145 can maintain the candidate language 235' from the candidate set 240''. Otherwise, when the candidate language 235' is determined not to be found in any of the information resources 420 of the query results 415, the search evaluator 145 can remove the candidate language 235' from the candidate set 240.
[0065] For each candidate language 235'' identified from the information resources 420, the search evaluator 145 can compute, determine, or otherwise generate a confidence score. The confidence score can indicate a probability or degree of certainty that the user 165 actually uses the corresponding candidate language 235''. In the computation, the search evaluator 145 can identify a number of occurrences of the candidate language 235'' from the information resources 420 of the query results 415. In some implementations, the search evaluator 145 can maintain a counter to track the number of occurrences of the candidate language 235'' identified from parsing the information resources 420 of the query results 415. Further, the search evaluator 145 can identify one or more orders of the information resources 420 identified as the candidate language 235'' from the rankings 425. As discussed above, the rankings 425 can indicate a degree of relevance of the information resources 420 to the keyword 230 and can identify an order of the information resources 420 in the query results 415.
[0066] Based on the order and the number of occurrences identified from the ranking 425 of the information resources 420, the search evaluator 145 can generate a confidence score for each candidate language 235". In some implementations, the search evaluator 145 can use the frequency of occurrence of the corresponding language 235" to determine the confidence score. The frequency can be based on the number of occurrences of the corresponding candidate language 235" and the total number of occurrences of all identified candidate languages 235". Generally, the higher the number of occurrences and the higher the order in the ranking 425, the higher the confidence score of the candidate language 235" can be. Conversely, the lower the number of occurrences and the lower the order in the ranking 425, the lower the confidence score of the corresponding candidate language 235" can be.
[0067] Using the confidence scores, the search evaluator 145 can determine whether to add or include the candidate language 235" in the candidate set 240". In some implementations, the search evaluator 145 can select the candidate languages 235" that correspond to the N highest confidence scores to include in the candidate set 240". In some implementations, the search evaluator 145 can compare the confidence score of the corresponding candidate language 235" to a threshold score to determine whether to include in the candidate set 240". The threshold score can delineate or divide the value of the confidence score of the corresponding candidate language 235" that will be included in the candidate set 240". When the confidence score meets (e.g., is greater than or equal to) the threshold score, the search evaluator 145 can select the corresponding candidate language 235" to include in the candidate set 240". On the other hand, when the confidence score does not meet (e.g., is less than) the threshold score, the search evaluator 145 does not select the corresponding candidate language 235" to include in the candidate set 240".
[0068] As already described above, by using the confidence score to determine whether to include a candidate language in the candidate set, the number of different languages in the candidate set can only increase when the candidate language has been sufficiently used (i.e., meets or exceeds the threshold score). Thus, this avoids increasing the size of the candidate set, and thus, the number of content items that can be identified by the method. By accurately limiting the number of content items to only those languages that the user can use, the method can more efficiently provide content items in a language that the user can understand.
[0069] Referring now to Figure 5A sequence diagram depicting a content selection process 500 of the system 100 for automatically detecting a user's language for content selection is illustrated. As shown, the language evaluator 150 running on the data processing system 110 can determine or identify one or more languages (e.g., languages 165A and 165B) of a language set 505 used by the user 160 from the candidate languages 235, 235', 235" of the candidate sets 240, 240', 240". In some implementations, the language evaluator 150 can omit the candidate set 240 (and candidate languages 235) from this determination. In some implementations, the language evaluator 150 can determine or identify the intersection between the candidate sets 240, 240', 240" to identify common candidate languages 235, 235', 235". When one or more candidate languages 235, 235', 235" are found in all of the candidate sets 240, 240', 240", the language evaluator 150 can identify or determine them as common. In contrast, when one or more candidate languages 235, 235', 235" are found in less than all of the candidate sets 240, 240', 240", the language evaluator 150 can identify or determine them as not common. Based on this intersection, the language evaluator 150 can determine or identify the common candidate languages 235, 235', 235" as the languages used by the user 160 for the language set 505.
[0070] The language evaluator 150 can associate the identified languages of the language set 505 (e.g., the depicted languages 165A and 165B) with the account profile 220. The language evaluator 150 can also store and maintain the association of the account profile 220 with the one or more languages of the language set 505 to the database 300. The association can be in the manner of one or more data structures (e.g., linked lists, arrays, trees, entries on a DMBS) stored and maintained on the database 300. Conversely, the language evaluator 150 can also determine or identify the candidate languages 235, 235', 235" outside of the intersection between the candidate sets 240, 240', 240" as not used by the user 160 associated with the client device 125. In some implementations, the language evaluator 150 can identify the languages outside of the intersection as not associated with the account profile 220. The language evaluator 150 can also store and maintain the lack of association of the account profile 220 to the database 300. The lack of association can be in the manner of one or more data structures (e.g., linked lists, arrays, trees, entries on a DMBS) stored and maintained on the database 300.
[0071] The content aggregator 155 running on the data processing system 110 can maintain a collection of content items 510 from one or more content providers 115 on the database 300 (or a separate database). Each content item 510 can correspond to or include text, image, audio, video, or multimedia content to be presented via a client device 125. The content item 510 can correspond to or include an object to be inserted onto an information resource (e.g., the information resource 210). According to HTML5, the object can be, for example, an iframe, a text object, an image, an audio object, a canvas object, or a video object, among others. Each content item 510 can be referenced by an identifier, such as a URL or another set of alphanumeric characters, among others.
[0072] In some embodiments, the content aggregator 155 can retrieve, identify, or receive the content items 510 themselves from the content providers 115 via the network 105. Upon receipt, the content aggregator 155 can store and maintain the content items 510 on the database 300. In some embodiments, the content aggregator 155 can retrieve, identify, or receive identifiers of the content items 510 from the content providers 115. The identifiers of the content items 510 can reference or correspond to locations of the content items 510 stored or maintained by the content providers 115, and can be, for example, a URL or another set of alphanumeric characters, among others. Upon receipt, the content aggregator 155 can store and maintain the identifiers of the content items 510 on the database 300.
[0073] The content items 510 can include content in one or more languages 165 (e.g., the depicted first language 165A and second language 165B). For example, as depicted, the content items 510 can include content items 510A-1 through 510A-X in the first language 165A (hereinafter collectively referred to as content items 510A). The content items 510 can also include content items 510B-1 through 510B-X in the second language 165B (hereinafter collectively referred to as content items 510B). Each content item 510 can be associated with at least one selection criterion. The selection criterion can specify, define, or identify parameters according to which the associated content item 510 will be selected as a candidate for provision to the client device 125. For example, the content items 510 can include text and images of a soccer ball provided by "XYZ" company. In this example, the associated selection criterion 510 can specify that the client device 125 previously visited an information resource (e.g., a web page) containing content related to the soccer ball or the company. The parameters of the selection criterion can include account segment, geographic region, and device type, among others. The selection criterion can be configured or set by the content provider 115 providing the content items 510 to the data processing system 110.
[0074] In some implementations, the content item 510 is identified as being available in one language by the content provider 115. For example, when submitting the content item 510 to the data processing system 110, the content provider 115 can send an indication of the language 165 (e.g., as one of the first language 165A or the second language 165B) that labels the content item 510. In some implementations, the content item 510 is identified as being available in one language 165 by the language evaluator 150 in the manner described above. For example, upon receiving the content item 510, the content aggregator 155 can apply the language identification model 305 to the content of the content item 510 to determine the language of the content item 510.
[0075] In some implementations, the content aggregator 155 can verify or determine that the language of the content item 510 is the same language as the associated information resource. The information resource can be associated via a link contained in the content item 510. For example, the associated information resource can be a landing page of the content item 510. To verify, the content aggregator 155 can identify the information resource associated with the content item 510 (e.g., via the link). The content aggregator 155 can compare the language used by the content item 510 to the language used by the associated information resource. The content aggregator 155 can determine the language of the content item 510 by applying the language identification model 245 to the content item 510. Further, the content aggregator 155 can determine the language of the associated information resource by applying the language identification model 245 to the information resource. When the languages are determined to match or correspond, the content aggregator 155 can include or add the content item 510 to the candidate set for the respective language. Otherwise, when the languages are determined to not match or correspond, the content aggregator 155 can exclude the content item 510 from the candidate set for the respective language.
[0076] Reference is now made to Figure 6FIG. 6B depicts a sequence diagram of a results provision process 600 for the system 100 for automatically detecting a user language for content selection. As shown, the content aggregator 155 can identify or select at least one content item 510' to provide to the client device 125. The selection of the content item 510' can be from a set of content items 510A in a first language 165A and a set of content items 510B in a second language 165B. In some implementations, the content aggregator 155 can generate, determine, or identify a selection value for each identified content item 510. The selection value can be used to identify the at least one content item 510' to provide to the client device 125 for presentation. The determination of the selection value for the content item 510 can be based on a comparison between the request 225 and selection criteria for the content item 510. For example, the content aggregator 155 can determine the selection value by comparing the keywords 230 in the request 225, the segments of the account profile 202, and the device type and location of the client device 125, among others, to the selection criteria for the content item 510 that determines the selection value.
[0077] Using the selection values for the content items 510, the content aggregator 155 can select the content item 510' from the set of content items 510A in the first language 165A and the set of content items 510B in the second language 165B. In some implementations, the content aggregator 155 can select the content item 510' that corresponds to the highest selection value. In some implementations, the content aggregator 155 can select the content item 510' according to a content selection protocol. The content selection protocol can include, for example, a real-time bidding protocol and a header bidding protocol, among others. The operation of the content selection protocol can be distributed among the data processing system 110, the content providers 115, and the client device 125. In performing the content selection protocol, the content aggregator 155 can retrieve, identify, or receive a bid value (e.g., a bid price) from each content provider 115 that has a content item 510 in the candidate set 515A or 515B. In some implementations, the content aggregator 155 can combine the bid value with the selection value for the content item 510 of the content provider 115 to modify or determine the selection value. In combining, the content aggregator 155 can identify or select the content item 510 that corresponds to the highest selection value to use as the selected content item 510'. The selected content item 510' can be from the candidate set in the first language or the candidate set in the second language.
[0078] With this selection, the content aggregator 155 can send, transmit, or provide the content item 510' to the client device 125. In some implementations, the content aggregator 155 can provide the information resource 420 (or an identifier of the information resource 420) identified from the search operation 405 to the content item 510'. The provision of the content item 510' and the information resource 420 can be via the at least one output 605. The application 205 can receive the content item 510' sent from the data processing system 110 via the network 105. Upon receipt, the application 205 can present the content item 510' on the information resource 215'. In some implementations, the application 205 can present the information resource 420 on the information resource 215' according to the ranking 425. For example, the information resource 215' can be a search results page, and can present corresponding identifiers of the information resource 420 along with the content item 510'.
[0079] In this way, the system 100 can improve the overall functionality of the data processing system 110 and the client device 125. By determining that the user 160 of the client device 125 is able to understand multiple languages 165A and 165B in a targeted manner, the candidate sets 515A and 515B can be expanded to include content items that use those languages 165A and 165B. Finally, the content item 510' selected from the candidate sets 515A and 515B can be in the language 165A or 165B, and can be provided for presentation to the user 160 operating the client device 125A. As a result, the information resource 215' can use the first language 165A, while the content item 510' inserted into the content slot 610 can use the second language 165B. By eliminating the necessity for separate queries for content using those languages 165, the inclusion of content using multiple languages 165A and 165B can reduce the consumption of computing resources at both the client device 125 and the data processing system 110. Furthermore, the human-computer interaction (HCI) between the user 160 and the system 100 can be enhanced by presenting content in potentially multiple languages 165.
[0080] Referring now to Figure 7 , a flow diagram of a method 700 of automatically detecting user language for content selection is depicted. The method 700 can be implemented using or performed by any of the components described in detail herein in connection with Figures 1 to 6 and Figure 8 The method 700 can also include any of the components described in detail herein in connection with Figures 1 to 6 and Figure 8The actions, operations, and functions of any of the components of the detailed description. Briefly, a data processing system can receive a request for content (705). The data processing system can determine a candidate language from the request for content (710). The data processing system can determine a candidate language from log records (715). The data processing system can determine a candidate language from search results (720). The data processing system can identify a language used (725). The data processing system can select a content item (730). The data processing system provides an output with the content item (735).
[0081] In more detail, a data processing system (e.g., data processing system 110) can receive a request for content (e.g., request 225) (705). The request for content can include one or more keywords (e.g., keywords 230) from a client device (e.g., client device 125). The keywords can be part of a search query and can be used to identify indexed information resources. The request can identify an account profile (e.g., account profile 220) or be associated with the account profile.
[0082] The data processing system can determine a candidate language (e.g., candidate language 235) from the request for content (710). The data processing system can parse the request to identify a language configuration of the client device or a language setting of the account profile. In addition, the data processing system can use a model (e.g., language identification model 245) to identify a language used by the keywords. From the parsing, the data processing system can identify a candidate language to include to a candidate set (e.g., candidate set 240).
[0083] The data processing system can determine a candidate language (e.g., candidate language 235') from log records (e.g., log records 305) (715). The data processing system can identify one or more activities maintained on log records of the client device or the account profile. For each identified activity, the data processing system can identify associated content. The data processing system can determine a language used by the content associated with the activity by applying a model. The data processing system can add the candidate language to the candidate set (e.g., candidate set 240').
[0084] The data processing system can determine a candidate language from search results (e.g., query results 415) (720). Using the keywords of the request for content, the data processing system can perform a search operation (e.g., search operation 405). From the search operation, the data processing system can identify one or more indexed information resources (e.g., information resources 420). The data processing system can apply a model to determine a language used by the information resources. The data processing system can add the candidate language to the candidate set (e.g., candidate set 240").
[0085] The data processing system can identify the language used (e.g., languages 165A and 165B) (725). The data processing system can determine an intersection between the candidate set of languages. The intersection can include one or more languages that are common across the candidate set. Using the intersection, the data processing system can identify the language used by the client device.
[0086] The data processing system can select a content item (e.g., content item 510') (730). The content item can be in one of the languages identified as used by the client device. The data processing system can identify the content item according to a content selection protocol. The data processing system can provide an output with the content item (e.g., output 605) (735). The output can include the selected content item, along with the indexed information resource.
[0087] Referring now to Figure 8 FIG. 8 shows a general architecture of an illustrative computer system 800 that can be used to implement any of the computer systems discussed herein, including the data processing system 110 and its components, the content provider 115, the content publisher 120, and the client device 125, in accordance with some embodiments. The computer system 800 can be used to provide information for display via a network 830. The computer system 800 includes one or more processors 820 communicatively coupled to a memory 825, one or more communication interfaces 805 communicatively coupled to at least one network 830 (e.g., the network 105), and one or more output devices 810 (e.g., one or more display units) and one or more input devices 815.
[0088] The processor 820 can include a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc., or combinations thereof. The memory can include, but is not limited to, electronic, optical, magnetic, or any other storage or transmission device capable of providing the processor with program instructions. The memory 825 can include any computer-readable storage media, and can store computer instructions, such as processor-executable instructions for implementing various functionality described herein for the various systems, as well as any data related thereto, resulting therefrom, or received via the communication interface or input devices (if present). The memory 825 can include a floppy disk, CD-ROM, DVD, magnetic disk, memory chip, ASIC, field-programmable gate array (FPGA), read-only memory (ROM), random-access memory (RAM), electrically erasable ROM (EEPROM), erasable- programmable ROM (EPROM), FLASH memory, optical media, or any other suitable memory from which the processor can read instructions. The instructions can include code from any suitable computer programming language.
[0089] Figure 8The processor 820 shown can be used to execute instructions stored in memory 825, and in doing so, can also read from or write to memory various information processed and / or generated according to the instructions. The processor 820 coupled to memory 825 (collectively referred to herein as a processing unit) can be included in components of system 100, such as data processing system 110 (and also content provider 115, content publisher 120, client device 125, and indexing service 130). For example, data processing system 110 may include memory 825 as a database. The processor 820 coupled to memory 825 (collectively referred to herein as a processing unit) can be included in content provider 115. For example, content provider 115 may include memory 825 to store content item 505 or 505'. The processor 820 coupled to memory 825 (collectively referred to herein as a processing unit) can be included in content publisher 120. For example, content publisher 120 may include memory 825 to store information resource 210. A processor 820 (collectively referred to herein as a processing unit) coupled to memory 825 may be included in client device 125.
[0090] The processor 820 of the computer system 800 can also be communicatively coupled to or configured as a control communication interface 805 to send or receive various information according to the execution of instructions. For example, the communication interface 805 can be coupled to a wired or wireless network, a bus, or other communication device, thus allowing the computer system 800 to send information to or receive information from other devices (e.g., other computer systems). Although not explicitly stated in Figures 1 to 6 While explicitly shown in the system, one or more communication interfaces facilitate the flow of information between components of system 800. In some implementations, the communication interfaces may be configured (e.g., via various hardware or software components) to provide a website as an access portal to at least some aspects of computer system 800. Examples of communication interfaces 805 include user interfaces (e.g., application 205, information resource 420 or 215', and content item 505 or 505') through which users can communicate with other devices of system 100.
[0091] Can provide Figure 8 The output device 810 of the computer system 800 shown may, for example, allow viewing or otherwise perceiving various information in conjunction with the execution of instructions. For example, an input device 815 may be provided to allow a user to manually adjust, make selections, input data, or interact with the processor in any of a variety of ways during the execution of instructions. Further information relating to the general computer system architecture applicable to the various systems discussed herein is provided herein.
[0092] The network 830 can include a computer network, such as the Internet, a local area network, a wide area network, a metropolitan area network, or other regional network, an intranet, a satellite network, other computer networks such as voice or data mobile telephony communication networks, and combinations thereof. The network 830 can be any form of computer network that relays information between components of the system 100, such as the data processing system 110 and its components, the content provider 115, the content publisher 120, the client device 125, and the index service 130. For example, the network 830 can include the Internet and / or other types of data networks, such as a local area network (LAN), a wide area network (WAN), a cellular network, a satellite network, or other types of data networks. The network 830 can also include any number of computing devices (e.g., computers, servers, routers, network switches, etc.) configured to receive and / or transmit data within the network 830. The network 830 can also include any number of hardwired and / or wireless connections. The client device 125 can be in wireless communication with a transceiver (e.g., via WiFi, cellular, radio, etc.) that is hardwired (e.g., via a fiber optic cable, a CAT5 cable, etc.) to other computing devices in the network 830.
[0093] Embodiments of the subject matter and operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by, or to control the operation of, data processing apparatus. The program instructions can be encoded on a propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The computer storage medium can be, or include, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, while a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium can also be, or include, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).
[0094] The features disclosed herein can be implemented in a smart TV module (or connected TV module, hybrid TV module, etc.), which may include a processing module configured to integrate an Internet connection with more traditional television program sources (e.g., receiving signals via cable, satellite, wireless, or other means). The smart TV module may be physically integrated into a television set or may include a separate device such as a set-top box, Blu-ray or other digital media player, game console, hotel TV system, or other complementary device. The smart TV module may be configured to allow viewers to search and find videos, movies, photos, and other content on the internet, local cable TV channels, satellite TV channels, or stored on a local hard drive. A set-top box (STB) or set-top unit (STU) may include an information device that may contain a tuner and connect to the television set and external signal sources to convert the signals into content, which is then displayed on the television screen or other display devices. The smart TV module may be configured to provide a main screen or top-level screen, including icons for multiple different applications, such as web browsers and multiple streaming services, connected cable or satellite media sources, other network "channels," etc. The smart TV module may also be configured to provide users with an electronic program guide. The accompanying application for the smart TV module can operate on a mobile computing device to provide the user with additional information about available programming, allowing the user to control the smart TV module, etc. In some embodiments, the features can be implemented on a laptop computer or other personal computer, smartphone, other mobile phone, handheld computer, tablet computer, or other computing device. In some embodiments, the features disclosed herein can be implemented on a wearable device or component (e.g., a smartwatch), which may include a processing module configured to integrate an Internet connection (e.g., with another computing device or network 830).
[0095] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or on data received from other sources. The terms "data processing apparatus," "data processing system," "user equipment," or "computing device" include all types of means, devices, and machines for processing data, including, for example, programmable processors, computers, one or more systems-on-a-chip, or combinations thereof. The apparatus may include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the apparatus may also include code that creates a runtime environment for the computer program in question, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. The apparatus and runtime environment can implement a variety of different computing model infrastructures, such as network services, distributed computing, and grid computing infrastructures.
[0096] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.
[0097] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit) and
[0098] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), for example. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0099] To provide for interaction with a user, implementations of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), plasma, or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can include any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s client device in response to requests received from the web browser.
[0100] Implementations of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), an inter-network (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0101] A computing system such as system 800 or system 100 can include servers and clients. For example, data processing system 110 of system 100 and its components, content provider 115, content publisher 120, client device 125, and index service 130 can each include one or more servers in one or more data centers or server farms. A client (e.g., client device 125) and a server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., HTML pages) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., a result of the user interaction) can be received from the client device at the server.
[0102] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions, but rather as descriptions of features specific to particular implementations. Certain features that are described in this specification in the context of separate implementations can also be implemented in combination with each other. Conversely, various features that are described in the context of a single implementation can also be implemented separately or in any suitable subcombination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0103] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such an order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing can be advantageous. Moreover, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems can generally be integrated in a single software product or packaged into multiple software products. For example, the query processor 135, the profile exporter 140, the search evaluator 145, the language evaluator 150, and the content aggregator 155 can be part of the data processing system 110, a single module, a logical device with one or more processing modules, or one or more servers.
[0104] In some cases, multitasking and parallel processing can be advantageous. Moreover, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems can generally be integrated in a single software product or packaged into multiple software products. For example, the query processor 135, the profile exporter 140, the search evaluator 145, the language evaluator 150, and the content aggregator 155 can be part of the data processing system 110, a single module, a logical device with one or more processing modules, or one or more servers.
[0105] In situations in which the systems discussed here collect personal information about users, or can make use of personal information, the users can be provided with an opportunity to control whether programs or features collect user information (e.g., information about a user's social network, social actions or activities, a user's preferences, or a user's location), or to control whether and / or how to receive content from a content server or other data processing system. In addition, certain data can be anonymized in one or more ways before it is stored or used, so that personally identifiable information is removed when generating parameters. For example, a user's identity can be anonymized so that no personally identifiable information can be determined for the user, or a user's geographic location can be generalized where location information is obtained (such as to a city, postal code, or state level), so that a particular location of a user cannot be determined. Thus, the user can have control over how information is collected about him or her and used by a content server.
[0106] Some illustrative implementations have now been described, and it is clear that the foregoing is illustrative rather than limiting, and has been presented by way of example. In particular, while many of the examples presented herein involve specific combinations of method actions or system elements, these actions and those elements can be combined in other ways to achieve the same objective. The behaviors, elements, and features discussed in connection with only one implementation are not intended to exclude similar roles from other implementations or implementations.
[0107] The wording and terminology used herein are for descriptive purposes and should not be considered limiting. The terms “comprising,” “including,” “having,” “containing,” “involving,” “characterized as,” “featured as,” “characterized in,” and variations thereof, as used herein, are intended to include items listed thereafter, their equivalents, and additional items, as well as alternative implementations consisting of items specifically listed thereafter. In one implementation, the systems and methods described herein consist of more than one or all of each combination of the described elements, actions, or components.
[0108] Any reference to an implementation, element, or action of a system or method mentioned herein in the singular may also include implementations that include multiple such elements, and any plural reference to any implementation, element, or action herein may also include implementations that include only a single element. References in the singular or plural form are not intended to limit the currently disclosed systems or methods, their components, actions, or elements to a singular or plural configuration. A reference to any action or element based on any information, action, or element may include an implementation in which the action or element is at least partially based on any information, action, or element.
[0109] Any implementation disclosed herein may be combined with any other implementation, and references to “an implementation,” “some implementations,” “alternative implementations,” “various implementations,” “one implementation,” etc., are not necessarily exclusive to each other and are intended to indicate that a particular feature, structure, or characteristic described in connection with an implementation may be included in at least one implementation. These terms used herein do not necessarily all refer to the same implementation. Any implementation may be combined inclusively or exclusively with any other implementation in any manner consistent with the aspects and implementations disclosed herein.
[0110] A reference to "or" can be interpreted as inclusive, such that any term described using "or" can refer to a single, more than one, or any of the terms described.
[0111] Where a technical feature is followed in the drawings, the detailed description or any claim by a reference sign, the sole purpose of the reference sign is to increase the intelligibility of the drawings, detailed description and claims. Therefore, whether a reference sign or its absence has no limiting effect on the scope of any claim element.
[0112] The systems and methods described herein can be embodied in other specific forms without departing from the characteristics thereof. Although the examples provided herein relate to selecting content to be provided in a networked environment, the systems and methods described herein can include application to other environments. The foregoing implementation is illustrative rather than limiting, and the described systems and methods are not limited to the foregoing implementation. The scope of the systems and methods described herein is thus indicated by the appended claims rather than by the foregoing description, and changes that come within the meaning and range of equivalency of the claims are embraced therein.
Claims
1. A computer-implemented method, comprising: A data processing system having one or more processors receives a request from a client device for the contents of an account profile. The data processing system uses browsing history logs of identified account profiles to analyze the logs to determine a first candidate language set from multiple languages, wherein the language recognition model is trained on a training dataset of a corpus containing text from each of the multiple languages in the following manner: By employing the language recognition model, each of the text corpora in each of the multiple languages is applied to generate a result set corresponding to the multiple languages. The result error is generated by comparing each of the resulting languages with the marked languages of each of the corpus. Based on the error in the result, one or more weights of the language recognition model are modified; The data processing system determines a second set of candidate languages from multiple languages based on the language settings in the account profile. The data processing system identifies a set of languages included in both a first candidate language set and a second candidate language set, the language set comprising a first language and a second language from multiple languages; and The data processing system stores account profiles and the association between the first and second languages in one or more data structures.
2. The method according to claim 1, further comprising: The data processing system generates a first confidence score for the first language based on the first occurrence of the first language in at least one of the language settings in the browsing history or account profile. and The data processing system generates a second confidence score for the second language based on the number of times the second language appears in at least one of the language settings in the browsing history or account profile.
3. The method according to claim 1, further comprising: In response to determining that the first confidence score of the first language is greater than the threshold score, the data processing system includes the first language in the first candidate language set; and In response to the data processing system determining that the second confidence score of the second language is greater than the threshold score, the second language is included in the second candidate language set.
4. The method according to claim 1, further comprising: The data processing system identifies a first plurality of content items in the first language and a second plurality of content items in the second language; and The data processing system provides a selected content item from a first plurality of content items and a second plurality of content items to the client device, the content item being in a first language or a second language.
5. The method according to claim 1, further comprising: The data processing system identifies the selection value of each content item in the first plurality of content items in the first language and the second plurality of content items in the second language; and The data processing system selects content items from a first plurality of content items and a second plurality of content items according to a content selection protocol to provide to the client device, the content items being in one of a first language or a second language.
6. The method according to claim 1, further comprising: Information resources that are identified by the data processing system as being associated with content items in the first or second language; The data processing system determines that the language of the content item corresponds to the language of the information resource; and In response to determining that the language of a content item corresponds to the language of the information resource, the data processing system adds the content item to a plurality of candidate content items for the client device to select from.
7. The method according to claim 1, further comprising: The data processing system identifies a third candidate language set from at least one of the following: (i) the content in each of a plurality of information resources identified in response to a request for content and the corresponding ranking of each information resource; (ii) the language configuration of an application running on a client device; or (iii) one or more keywords included in the request for content. and The identification of the language set also includes the identification of the language set included in the first candidate language set, the second candidate language set, and the third candidate language set.
8. The method according to claim 1, wherein, Determining the first candidate language set also includes determining from multiple languages based on browsing history identified in log records, the browsing history including at least one of the following: search queries received from the client device, access to information resources by the client device, and interactions with elements on the information resources.
9. The method according to claim 7, wherein, Determining the third candidate language set also includes: identifying from multiple languages the frequency of each language in the third candidate language set across multiple information resources, based on responses to requests for content.
10. The method according to claim 1, wherein, Receiving a request for content also includes receiving a query that includes one or more keywords; and further includes: Perform a search using one or more keywords in the query to identify multiple information resources; and The data processing system provides the client with an output comprising at least one of a plurality of information resources and a selected content item from a first plurality of content items in a first language and a second plurality of content items in a second language, wherein the content item is in one of the first or second languages.
11. A computer-implemented system, comprising: A data processing system, having one or more processors coupled to memory, is configured as follows: Receive a request from the client device to identify the contents of the account profile; Using browsing history logs of identified account profiles, a first candidate language set is determined from multiple languages by analyzing the logs using a language recognition model, wherein the language recognition model is trained on a training dataset based on a corpus containing text from each of the multiple languages in the following manner: By employing the language recognition model, each of the text corpora in each of the multiple languages is applied to generate a result set corresponding to the multiple languages. The result error is generated by comparing each of the resulting languages with the marked languages of each of the corpus. Based on the error in the result, one or more weights of the language recognition model are modified; Based on the language settings in the account profile, a second set of candidate languages is determined from multiple languages; Identify a language set included in both a first candidate language set and a second candidate language set, the language set comprising a first language and a second language from a variety of languages; and Store account profiles and associations between the first and second languages using one or more data structures.
12. The system according to claim 11, wherein, The data processing system is also configured to: A first confidence score for the first language is generated based on the number of times the first language appears in at least one of the language settings in the browsing history or account profile. and A second confidence score for the second language is generated based on the number of times the second language appears in at least one of the language settings in the browsing history or account profile.
13. The system according to claim 11, wherein, The data processing system is also configured as follows: In response to determining that the first confidence score of the first language is greater than the threshold score, the first language is included in the first candidate language set; and In response to determining that the second confidence score of the second language is greater than the threshold score, the second language is included in the second candidate language set.
14. The system according to claim 11, wherein, The data processing system is also configured as follows: Identify the first multiple content items in the first language and the second multiple content items in the second language; and Provide a client device with a content item selected from a first plurality of content items and a second plurality of content items, the content item being in a first language or a second language.
15. The system according to claim 11, wherein, The data processing system is also configured as follows: Identify the selection value of each content item in the first plurality of content items in the first language and the second plurality of content items in the second language; and According to the content selection protocol, content items are selected from a first plurality of content items and a second plurality of content items to be provided to the client device, the content items being in one of a first language or a second language.
16. The system according to claim 11, wherein, The data processing system is also configured as follows: Identify information resources associated with content items in the first or second language; The language of the content item corresponds to the language of the information resource; and In response to determining that the language of the content item corresponds to the language of the information resource, the content item is added to a plurality of candidate content items for the client device to select from.
17. The system according to claim 11, wherein, The data processing system is also configured as follows: A third candidate language set is identified from at least one of the following: (i) the content in each of a plurality of information resources identified in response to a request for content and the corresponding ranking of each information resource; (ii) the language configuration of an application running on a client device; or (iii) one or more keywords included in the request for content. and Identify the language sets included in the first candidate language set, the second candidate language set, and the third candidate language set.
18. The system according to claim 11, wherein, The data processing system is also configured to determine from multiple languages based on browsing history identified in log records, the browsing history including at least one of the following: search queries received from client devices, access to information resources by client devices, and interactions with elements on information resources.
19. The system according to claim 17, wherein, The data processing system is also configured to determine the frequency of each language from a third set of candidate languages across multiple information resources, identified in response to requests for content.
20. The system according to claim 11, wherein, The data processing system is also configured as follows: Receive a query from the client device that includes one or more keywords; and Perform a search using one or more keywords in the query to identify multiple information resources; The output provided to the client device includes at least one of a plurality of information resources and a content item selected from a first plurality of content items in a first language and a second plurality of content items in a second language, the content item being in one of the first language or the second language.
Citation Information
Patent Citations
Automatically determining language for speech recognition of spoken utterance received via an automated assistant interface
CN110998717A
Providing language recommendations
US20150347378A1
Interaction context-based virtual reality
US20180095624A1
Language-specific search results
US8375025B1
Language-specific search results
US9275113B1