All articles
Formatters
7 min readBy DevUtilX Team

XML Formatting: Structure, Namespaces and Pretty-Printing

Learn how XML is structured, what makes it well-formed, how namespaces work, and how to pretty-print XML safely for debugging and review.

XML Formatting: Structure, Namespaces and Pretty-Printing

XML has a reputation as the older, heavier cousin of JSON, but it has not gone away. It still carries enterprise integrations, office documents, Android layouts, SVG graphics, RSS and Atom feeds, SOAP services, sitemaps and build configuration. If you work on any of these, you will read XML, and you will often meet it as one long unreadable line. This guide explains how XML is structured, what makes a document well-formed, how namespaces work, and how to pretty-print XML so it can be read, reviewed and debugged.

What is XML?

XML, the Extensible Markup Language, is a text format for representing structured data as a tree of named elements. It is defined by the W3C in the XML 1.0 specification, and it was designed to be both human-readable and machine-readable, with a strict grammar so that every parser behaves the same way.

The word "extensible" matters. Unlike HTML, which has a fixed set of tags, XML lets you invent your own vocabulary. A bookstore, a payroll system and a vector graphic can all use XML, each with different element names. The MDN XML introduction gives a gentle overview.

A small document looks like this:

<?xml version="1.0" encoding="UTF-8"?>
<library>
  <book id="b1" available="true">
    <title>Learning XML</title>
    <author>Ada Lovelace</author>
    <price currency="USD">29.99</price>
  </book>
</library>

The building blocks

XML has a small set of constructs. Knowing each one makes any document easier to read.

  • XML declaration. The optional first line, <?xml version="1.0" encoding="UTF-8"?>, states the version and character encoding. If present, it must be the very first thing in the file.
  • Elements. The main unit, written as a start tag, content and an end tag: <title>Learning XML</title>. An empty element can be written as <br/>.
  • Attributes. Name and value pairs inside the start tag, such as id="b1". Values must always be quoted.
  • Text content. Character data between tags.
  • Comments. Written <!-- like this -->.
  • CDATA sections. Blocks written <![CDATA[ ... ]]> in which characters such as < and & are treated literally.
  • Processing instructions. Directives for applications, such as a stylesheet link.
  • Entity references. Five predefined escapes: &lt;, &gt;, &amp;, &quot; and &apos;.

Well-formed versus valid

Two terms are often mixed up.

A document is well-formed when it obeys XML's syntax rules. A parser must reject a document that is not well-formed, and that is the key difference from HTML, which browsers repair silently. The rules are:

  1. There is exactly one root element.
  2. Every start tag has a matching end tag, or is self-closing.
  3. Elements are properly nested: <a><b></b></a>, never <a><b></a></b>.
  4. Attribute values are quoted, and an attribute name appears only once per element.
  5. Names are case-sensitive: <Item> and <item> are different.
  6. Special characters in text and attributes are escaped.

A document is valid when it is well-formed and also conforms to a schema that describes the allowed elements, order and types. Schemas can be written as a DTD, XML Schema (XSD), or RELAX NG. Well-formedness is about syntax. Validity is about meaning.

Why XML formatting matters

Whitespace between elements is usually insignificant, so these two documents describe the same data:

<order id="1"><item sku="A1" qty="2"/><item sku="B7" qty="1"/></order>
<order id="1">
  <item sku="A1" qty="2"/>
  <item sku="B7" qty="1"/>
</order>

People care about the difference:

  • Readability. Nested structure is obvious when each level is indented.
  • Debugging. SOAP responses, feed output and log payloads often arrive as a single line. Pretty-printing is the first step to finding the field that is wrong.
  • Diffs and reviews. One element per line means a version-control diff shows exactly what changed.
  • Spotting errors. Mismatched tags and misplaced attributes stand out in a regular layout.

A caution about whitespace

XML does not always ignore whitespace. In mixed content, where text and elements are interleaved, spaces can be part of the data:

<p>This is <b>important</b> text.</p>

If a formatter breaks that line after <b>, it can add or remove a visible space. Pretty-printing is safe for data-oriented XML (records, configuration, feeds). Treat it with care for document-oriented XML (prose with inline markup). Some formats also define their own whitespace rules, controlled by the xml:space attribute.

Namespaces

Namespaces solve a naming problem: what happens when two vocabularies use the same element name? A <title> in a book is not the same as a <title> in an HTML page. The Namespaces in XML recommendation lets each name belong to a vocabulary identified by a URI.

<feed xmlns="http://www.w3.org/2005/Atom"
      xmlns:media="http://search.yahoo.com/mrss/">
  <title>Example Feed</title>
  <entry>
    <title>First post</title>
    <media:thumbnail url="https://example.com/a.jpg"/>
  </entry>
</feed>

Key points:

  • xmlns="..." sets the default namespace for the element and its descendants.
  • xmlns:media="..." declares a prefix, used as media:thumbnail.
  • The URI is an identifier, not a link. It does not need to resolve to a web page.
  • The prefix is only a local alias. Two documents can use different prefixes for the same namespace and still mean the same thing.

Namespaces are the most common source of confusion when querying XML. An XPath expression such as //title finds nothing in the feed above unless you register the namespace and use a prefix in the query.

Real-world places you will meet XML

  • Web feeds: RSS and Atom.
  • Graphics: SVG is an XML vocabulary.
  • Sitemaps: the sitemap.xml that search engines read.
  • Office files: .docx, .xlsx and .pptx are zipped bundles of XML parts.
  • Android: layouts, manifests and resources.
  • Build and configuration tools: Maven's pom.xml, .NET project files and many IDE settings.
  • Enterprise integration: SOAP web services, banking and healthcare message formats.

Each of these is easier to work with after pretty-printing.

Common mistakes

  • Unescaped special characters. A bare & or < in text makes the document not well-formed. Use &amp; and &lt;, or a CDATA section.
  • Mismatched or misnested tags. Often caused by hand-editing or string concatenation.
  • Multiple root elements. XML allows only one.
  • Wrong or missing encoding. If the declaration says UTF-8 but the file is saved in another encoding, accented characters turn into garbage.
  • Case errors. <Name> is not closed by </name>.
  • Unquoted attributes. Valid in old HTML, a fatal error in XML.
  • Forgetting namespaces in queries. Your XPath is correct but returns nothing.
  • Building XML by string concatenation. This invites escaping bugs and injection problems. Use an XML library or serializer to generate output.
  • Parsing untrusted XML with defaults. External entities can let an attacker read local files. The OWASP XXE prevention cheat sheet explains how to disable the risky features.

Step-by-step: pretty-print XML with DevUtilX

For a quick clean-up without installing anything, use the DevUtilX XML Beautifier, which runs in your browser.

  1. Open the XML beautifier tool.
  2. Paste your XML into the input editor.
  3. Choose your indentation, such as two spaces or tabs.
  4. Run the formatter and review the indented output.
  5. Copy the result into your editor, ticket or documentation.

Your data is processed locally and not uploaded, which is useful when a payload contains customer records or internal identifiers. If the formatter reports an error, the document is probably not well-formed, so look for an unclosed tag, an unescaped & or a stray character before the declaration. The XML Validator can help pinpoint the problem. When you need the opposite, use the XML Minifier to strip whitespace for transport.

Working with XML in code

Pretty-printing also exists in most languages, which is useful for logs and tests.

In JavaScript, the browser provides DOMParser and XMLSerializer:

const doc = new DOMParser().parseFromString(xmlText, "application/xml");
const error = doc.querySelector("parsererror");
if (error) throw new Error(error.textContent);
console.log(new XMLSerializer().serializeToString(doc));

In Python, the standard library can indent a parsed tree:

import xml.etree.ElementTree as ET

tree = ET.ElementTree(ET.fromstring(xml_text))
ET.indent(tree, space="  ")
tree.write("out.xml", encoding="utf-8", xml_declaration=True)

Both approaches parse first, so they also confirm that the document is well-formed. For large files, prefer a streaming parser to avoid loading everything into memory.

Best practices

  • Generate XML with a library, never by concatenating strings.
  • Declare the encoding and save the file in that encoding. UTF-8 is the safe default.
  • Indent consistently and use one element per line for data-style documents.
  • Be careful with mixed content. Do not reformat prose that contains inline tags unless you know the whitespace is irrelevant.
  • Use meaningful, consistent names. Pick one naming style, such as camelCase or kebab-case, and keep it.
  • Prefer elements for data and attributes for metadata, and be consistent about the choice.
  • Validate against a schema at system boundaries so bad data is caught early.
  • Disable external entities and DTD processing when parsing untrusted input.
  • Keep formatted source in version control, and minify only for transmission or storage when size matters.

XML compared with JSON and YAML

Aspect XML JSON YAML
Readability Verbose but explicit Compact Very compact
Comments Yes No Yes
Namespaces and schemas Mature (XSD, DTD) JSON Schema Limited
Mixed content (text and markup) Yes No No
Typical use Documents, enterprise, feeds Web APIs Configuration

JSON is usually a better fit for web APIs because it is lighter and maps directly onto JavaScript objects. XML remains the right choice where you need mixed content, namespaces, mature schema tooling or compatibility with an existing standard. The DevUtilX XML to JSON and JSON to XML tools help when you must bridge the two, and our JSON Formatting Guide covers the other side.

FAQ

Is XML dead?
No. It is less common in new web APIs, but it underpins many standards and enterprise systems, along with office formats, feeds, SVG and Android.

What is the difference between well-formed and valid XML?
Well-formed means the syntax is correct. Valid means the document is well-formed and also follows a schema. Every valid document is well-formed, but not the other way around.

Does indentation change the meaning of XML?
Usually not in data-oriented documents. In mixed content, or where xml:space is set, whitespace can matter, so review the result.

Should I use attributes or child elements?
Use elements for data and attributes for metadata about that data, such as ids and units. The most important thing is to stay consistent.

Why does my XPath query return nothing?
The most common cause is a default namespace. Register the namespace with a prefix in your XPath engine and use that prefix in the expression.

Can I put comments in XML?
Yes, with <!-- -->. Comments cannot be nested, and they should not contain a double hyphen.

Is it safe to paste XML into an online formatter?
Only if the tool runs entirely in your browser. DevUtilX processes your input locally, but always check a tool's privacy statement first.

Conclusion

XML rewards care. A strict grammar, clear namespaces and consistent formatting make documents that people and programs can both trust. Learn the well-formed rules, use namespaces deliberately, generate XML with libraries rather than string concatenation, and pretty-print anything you need to read. For a quick clean-up of any document, try the XML Beautifier. To keep going, see our guides on HTML formatting and SQL formatting.

Further reading

Try the tools

Further reading

Related articles