Text Analytics supported languages

The following language requirements must be met to implement Medallia Text Analytics for a customer:

  1. The language is supported.
  2. There is a strategy to build rules for the language. For example, there is a fluent speaker of the language at Medallia, or a person at the company that speaks the language and will receive Text Analytics training.

In addition, if the company wants sentiment, verify that there is either a sentiment model available for the applicable languages, or that the languages use sentiment processing.

To request a new language for Text Analytics, contact your Medallia Expert.

If you have access to help.medallia.com, use the Product Enhancement Request form to request a new language.

Text Analytics language and sentiment availability

The technical components of Text Analytics language processing include parsing and sentiment. Parsing is the ability to ingest data and enable for rule-based topic building and theme discovery. Sentiment analysis identifies the level positivity of phrases in customer feedback. For more information about sentiment, see Sentiment analysis.

Text Analytics uses a tiered approach to describe which languages are supported and the functionality available in each language, based on whether languages support parsing and sentiment (Tier 1), or only parsing (Tier 2).

Tier 1 languages

The following languages support parsing and sentiment.

  • Arabic
  • Czech
  • Chinese (Traditional and Simplified)
  • Danish
  • Dutch
  • English
  • French
  • German
  • Hebrew
  • Hungarian
  • Italian
  • Japanese
  • Korean
  • Polish
  • Portuguese
  • Romanian
  • Russian
  • Spanish
  • Swedish

Tier 2 languages

Tier 2 languages support parsing (topics and themes), but not sentiment. Contact your Medallia Expert before using a Tier 2 language. Text Analytics supports the following Tier 2 languages:

  • Norwegian
  • Turkish
  • Slovak

Tier 2 languages do not support sentiment analysis, and Text Analytics relies on sentiment tagging for some reporting filtering capabilities. As a result, Text Analytics and Reporting functionality is limited for Tier 2 languages.

Reports can still be retrieved and topic alerts function as expected, but technical limitations include the following:

  • Phrases do not appear in reports filtered by topic.

  • R-fields cannot be used.

  • Some K-fields cannot be used.

  • Some Text Analytics reporting metrics cannot be used, including % Positive, % Negative, Impact Score, and Net Sentiment.

  • Topics are visible in reports and option boxes but cannot be viewed in detail.

  • Modules that show verbatims, like Responses Feed and Comment Stream, cannot be filtered by topic.

  • Phrase segmentation changes when sentiment is not available.

Supported languages for Action Intelligence features

The following table lists the languages Action Intelligence currently supports for various product functionality:

Product functionSupported languages
Attention scoringEnglish, French, German, Italian, Spanish
Suggested actionsChinese, English, French, German, Italian, Japanese, Portuguese, Spanish
EffortChinese, Czech, English, French, German, Italian, Japanese, Russian, Spanish
RecognitionChinese, Czech, English, French, German, Italian, Japanese, Portuguese, Russian, Spanish

If a text field contains a supported language, or has been translated to a supported language, Action Intelligence processes that field.

Note: Language processing for Action Intelligence and Text Analytics overlaps. Text Analytics supports all Tier 1 languages listed in Text Analytics language and sentiment availability, above. For each supported language you can configure whether to process comments in their native language or the translation, as described in Action Intelligence. If you choose to process the native language, Action Intelligence processes the same record only if the native language is English or Spanish. For example, if you configure Text Analytics to process Italian in the native language, Action Intelligence does not process those records.

Dialects

The Text Analytics parser and sentiment models are built to handle the traditional form of supported languages. (Chinese is an exception, as Text Analytics converts Traditional Chinese comments to the Simplified dialect, and parses Chinese using the Simple form of Mandarin.)

Dialect nuances are accounted for in the rule creation process. Medallia has native speakers available, who can ensure full coverage of dialect vocabulary and context nuances. For example, the English word train in European Portuguese is comboio, and in Brazilian Portuguese it is trem.

Vernacular

Use the following Text Analytics tools to handle geographical and sub-cultural vernacular/slang:

  • User features — Allow for the creation of a custom lexicon of words and phrases that might not be identified by a dictionary. This includes items such as shorthand used in contact center notes and various online chat platforms, emojis, names of company-specific products, services, rewards programs, and acronyms. For more information, see User features.
  • Word groups — Serve as a built in thesaurus that is editable by self-service Text Analytics users. Use word groups to create groupings of similar words, making topic building easier and more scalable. For more information, see Word groups for standard topics.
  • Topic rules — After creating user features and word groups, use topic rule syntax to capture true customer feedback, using the exact language used by customers. Because the Text Analytics parser examines raw language, you can make rules using whatever words your customers use. For more information, see Topics screen.

Machine Translations

Machine translation is not always the best approach for Text Analytics, and if a company is serious about getting insights from feedback in a certain language, it is not recommended. However, for certain languages (the ones that translate well with high BLEU scores) it can work well.

  • Western European languages are translated relatively well (German, French, Spanish, Italian, and so on)

  • Nordic languages are translated relatively well (Norwegian, Swedish, Danish, and so on)

  • Eastern European languages are not translated well (Russian, Polish, Czech, and so on)

If a language has a large enough volume for a company (~50,000 or more comments), always use the native language. However, if a company has a few primary languages with high volume and many low-volume languages, translation can be a good approach for those low-volume languages.

Important: Company stakeholders must understand that quality will be worse for any translated languages.  Medallia has no ability to modify or improve these translations.

Google

We currently support all publicly available Google languages for translation. For a full list see Google's Language support topic.

Important: With Google, you can translate from any supported language to one target language. For example, you can translate French, German, and Spanish to English or you can translate French, German, and English to Spanish.

Google is our highly preferred option. It is faster, more stable, and more accurate.

SYSTRAN

We currently support the translation of the following languages into English only.

Important: A customer should only use SYSTRAN if they cannot use Google for legal (Finance Vertical) or competitive reasons.  The quality of these translations is generally lower and the system is slower and less stable than Google.
  • Arabic

  • Chinese (Traditional and Simplified)

  • Czech

  • Danish

  • Dutch

  • Finnish

  • French

  • German

  • Greek

  • Hungarian

  • Italian

  • Japanese

  • Korean

  • Norwegian

  • Polish

  • Portuguese

  • Russian

  • Spanish

  • Swedish

  • Thai
  • Turkish