Go to main contentGo to footer
bazarjs
|
14 January 15

BazarJS: Module loaders and package managers

Second installment of BazarJS on the magical world of single-page applications... today we get serious with module loaders, bundlers and package managers, three of the most "confusing", controversial and peculiar topics in the JavaScript world.

As always, after an overview of the available solutions, we'll weigh the pros and cons and arrive at our own choice.

Stefano VernaHead of DatoCMS

As we all know, JavaScript has no native mechanism for managing dependencies between files: no equivalent, so to speak, of Ruby's require or Sass's @import. For years we got used to living without it by simply using a mix of anonymous functions and the global namespace:

// calculator.js
(function(root) {
  var calculator = {
    sum: function(a, b) { return a + b; }
  };
  root.Calculator = calculator;
})(this);

// app.js
console.log(calculator.sum(1, 2)); // => 3

Unfortunately, as we know, it's a "half" solution: since no explicit dependency tree is generated, developers are left to specify the order in which the different modules are included, which earns them the classic TypeError: undefined is not a function errors whenever they get the order wrong.

CommonJS

Node.js filled this gap by implementing the pattern specified by the grassroots committee called CommonJS. It's an elegant, convenient solution, similar to those found in other programming languages, and it provides the necessary module encapsulation:

// calculator.js
module.exports = {
  sum: function(a, b) { return a + b; }
};

// app.js
var assert = require('assert');
var calculator = require('./calculator');
assert.equal(3, calculator.sum(1, 2));

The require() function reads the content of the specified local file, evaluates it, and returns the content of the module.exports object to the caller. Anything in the module that isn't exposed on this object stays effectively "private" and inaccessible from outside.

Since the call to require() is synchronous, unfortunately it's a solution poorly suited to the browser, where dynamically loading JavaScript files must necessarily be asynchronous.

AMD

Precisely because of this, CommonJS defined an asynchronous variant for loading modules, called AMD (Asynchronous Module Definition), which is much easier to use in a browser:

// calculator.js
define("calculator", function() {
  return {
    sum: function(a, b) { return a + b; }
  };
});

// app.js
define("app", ["calculator"], function(calculator) {
  console.log(calculator.sum(1, 2)); // => 3
});

Modules are defined with the define() function, which also specifies the module name and any dependencies it needs. The pattern loads those dependencies asynchronously and then passes them to the callback as parameters, preserving the order specified in the array.

Less elegant than its Node.js equivalent, sure, but then again it's the only viable option in a browser (at least apparently... we'll come back to this later).

Light at the end of the tunnel: ECMAScript 6

ECMAScript 6, the next version of JavaScript, whose standardization was completed in 2014 and which browsers will slowly implement, has finally proposed an "official" solution to the problem, with a syntax similar to Python's, for example:

// calculator.js
export function sum(a, b) { return a + b; }

// app.js
import * as calculator from 'calculator';

console.log(calculator.sum(1, 2)); // => 3

Note that the import keyword is synchronous: does that mean it can only be used in Node.js and the like? Fortunately not: to implement this mechanism, browsers will have to statically analyze the code before actually evaluating the body of the JavaScript file, looking for any import statements to load ahead of evaluating the code itself.

And the developer? Pays the price.

Given today's fragmentation, and while waiting for universal support for the new "universal" ECMAScript 6 syntax, a JavaScript developer who wants to release a JavaScript library currently has to make their module compatible with all three loading options described above:

  • Via globals
  • CommonJS
  • AMD

It may look like a lot of effort, but in practice it's not particularly hard: you just wrap your library in a skeleton like this:

// calculator.js
(function (name, context, definition) {
  if (typeof module != 'undefined' && module.exports)
    module.exports = definition();
  else if (typeof define == 'function' && define.amd)
    define(name, definition);
  else
    context[name] = definition();
}('calculator', this, function () {
  // your module here!
  return {
    sum: function(a, b) { return a + b; }
  };
});

A library able to support all current loading "standards" is said to support UMD (Universal Module Definition).

Fortunately, most libraries developed in recent years support UMD, so they can be used in any runtime environment, both client-side and server-side.

But let's get down to business :)

After this necessary "theoretical" introduction, a question arises: which module-loading mechanism should we use to write our client-side application? Given the limits described, AMD would seem the only possible candidate... is that really so? And second: which module database should we draw from?

Let's explore the most popular solutions the JavaScript community has given us.

Solution 1: RequireJS + Bower

RequireJS is the most popular implementation of the AMD pattern around, while Bower is the go-to package manager for front-end packages (so not just JS, but also CSS, Sass, etc.).

Let's take a look at the numbers for these two projects:

##### RequireJS

* **Homepage:** http://requirejs.org/
* **Data di creazione:** Febbraio 2010
* **Github stars:** ★ 6.795
##### Bower

* **Homepage:** http://bower.io
* **Data di creazione:** Settembre 2012
* **Numero di moduli a disposizione:** 21.564
* **Github stars:** ★ 11.471

With RequireJS, dependencies are loaded on the page by importing a single script: RequireJS itself.

<script data-main="scripts/main" src="scripts/require.js"></script>

The data-main attribute on the tag specifies the application's entry point, which holds any RequireJS configuration needed to download the files, and which kicks off the chain of asynchronous AMD imports we've already seen:

// main.js
requirejs.config({ baseUrl: '/scripts' });
requirejs(['app']);

Since RequireJS only handles module loading, it clearly needs to be paired with a package manager like Bower, which gives us access to a huge database of third-party front-end modules/libraries (20,000+) and handles dependency analysis, any versioning conflicts, and the actual download of modules locally.

Bower's equivalent of the Gemfile is called bower.json, and it looks something like this:

{
  "name": "my-project",
  "private": true,
  "dependencies": {
    "rsvp": "~3.0.16"
  }
}

The bower install command reads this file and downloads the specified dependencies into the local ./bower_components directory.

Ça va sans dire, with the requirejs.config() method you can configure RequireJS to fetch Bower modules from this folder.

The harsh reality

AMD's asynchronous loading (and therefore RequireJS's) looks like a great idea at first glance, letting us progressively download only the files strictly needed to run our app.

In practice, it's an idea that's practically unworkable in a browser context, where the HTTP overhead of asynchronously downloading individual JavaScript files is dramatic, to the point of destroying application performance on moderately complex projects.

So RequireJS recommends running a bundling step on the modules with its command-line tool r.js:

node r.js -o name=main out=bundle.js baseUrl=.

The process parses the files starting from an entry point (in our case main.js) to rebuild the dependency tree based on the define() calls it finds along the way. It can then order and concatenate all the required modules into a single file, bundle.js, which we'll include in our HTML instead of the main.js file we saw earlier:

<script data-main="scripts/bundle" src="scripts/require.js"></script>

The bundle file is simply a concatenation of the JavaScript files needed at runtime:

// bundle.js
define("calculator", [],function() {
  return {
    sum: function(a, b) { return a + b; }
  };
});

define("app", ["calculator"], function(calculator) {
  console.log(calculator.sum(1, 2)); // => 3
});

requirejs(['app']);

The order in which the tool concatenates the files lets RequireJS cache the module bodies at runtime before reaching the initializing requirejs() call, thus avoiding any further download requests.

Solution 2: Browserify + Npm

Unlike RequireJS, Browserify lets you write your client-side application using synchronous CommonJS loading (à la Node.js, so to speak). How is that possible? We'll look at how in detail; meanwhile, a look at some usage stats never hurts:

##### Browserify

* **Homepage:** http://browserify.org/
* **Data di creazione:** Settembre 2010
* **Github stars:** ★ 6.164
##### Npm

* **Homepage:** http://npmjs.org/
* **Data di creazione:** Settembre 2009
* **Numero di pacchetti a disposizione**: 115.973 (!!!)
* **Github stars:** ★ 5.332

Like RequireJS, Browserify provides a command-line tool that parses modules starting from an entry point (in this case, app.js), building the tree of require() calls:

browserify app.js --outfile bundle.js

The result looks like this: [^debundle]

// bundle.js
debundle({
  entryPoint: "./app",
  modules: {
    "./app": function(require, module) {
      var calculator = require('./calculator');
      console.log(calculator.sum(1, 2));
    },
    "./calculator": function(require, module) {
      module.exports = {
        sum: function(a, b) { return a + b; }
      };
    }
  }
});

function debundle(data) {
  var cache = {};
  var require = function(name) {
    if (cache[name]) { return cache[name]; }
    var module = cache[name] = { exports: {} };
    data.modules[name](require, module);
    return module.exports;
  };
  return require(data.entryPoint);
}

You see what we did there? :) The starting entry point, along with the content of every module that might be required at runtime, is passed to a special debundle() function, which implements the simple CommonJS module.exports/require mechanism on the client side.

That's the whole trick: by bundling all the modules upfront, the require process can safely follow synchronous logic.

Not content with that, Browserify goes further: to fully simulate a Node.js environment, it lets you require() not only your own local application files, but also

  • any npm packages in the ./node_modules directory;
  • some of the Node.js core modules (url, path, stream, events, http).

Since these obviously weren't designed for the browser, the Browserify team rewrote them entirely while keeping the same API, and the bundling process includes them in place of the originals.

One last, important point: Browserify's parsing can be extended with so-called transforms, third-party logic that can modify/pre-process source files before they're included in the bundle:

browserify app.js         \
  --transform coffeeify   \
  --transform uglifify    \
  --outfile bundle.js

A command like this, for example, lets us write our app in CoffeeScript and get a compiled, minified bundle.js. Not bad.

[^debundle]: The debundle() function in this article has been simplified to make it easier to understand: here's the version Browserify actually uses.

Analysis

First of all, going back to the first part of this series, both solutions integrate well with the main task runners (e.g. gulp-browserify, gulp-requirejs), so it's 1-1 so far.

Browserify recognizes that RequireJS's asynchronicity is just a form of wishful thinking that's unworkable in practice [^http2], so it opts for the synchronous CommonJS pattern, which is more convenient and less verbose. A point in its favor.

[^http2]: Or at least it will be until HTTP/2 arrives, dramatically reducing latency and per-request overhead.

On the other hand, Browserify requires npm packages, when the package manager best suited to front-end use would actually be Bower. This is more a philosophical point than a practical one, given that:

  • the vast majority of front-end modules are also released on npm;
  • Browserify transforms like debowerify let you include Bower packages just like npm ones, if needed;

You still need to be careful about which npm packages you use, since they won't necessarily work in a browser. For exactly this purpose, Toby Ho released Browserify Search, a tool that gives us this information through fairly sophisticated analysis of npm packages. Fortunately, about half of npm packages (around 60,000) currently pass the checks.

Browserify currently has the largest and most active community among its competitors. Just to give an idea of how lively this project is:

  • Watchify, a watcher for Browserify that rebuilds the bundle whenever any project dependency is updated and, through caching, cuts build times after the first by an order of magnitude;
  • Disc, a tool that analyzes the Browserify bundle to show a navigable visual breakdown of the heaviest dependencies in the package;
  • partition-bundle, a plugin that splits the application's modules across multiple bundles, for a faster initial download and progressive loading of client-side logic.

Both RequireJS and Browserify support source maps, essential for debugging the application in production.

Our choice

Browserify. We chose this solution for the community support mentioned above and because it's more convenient to write, but that's not all.

Choosing to write a Node.js-compatible application brings two more huge advantages:

  • the ability to build isomorphic applications. In other words, JavaScript applications that, through simple abstractions at the routing and view-rendering level, can run the same way in the browser and on the server [^isomorphic]. This approach gets the best of both worlds:

    • extremely fast rendering of the first page the user loads, thanks to server-side pre-rendering, without waiting for all the JavaScript code to download and load in the browser;
    • instant client-side updates from the first load onward, keeping server calls to a minimum.
  • the ability to run unit tests on your front-end code without launching a browser, simply running them in a Node.js environment. We all know how essential it is to keep test run times as low as possible, and putting a browser in the middle forces a couple of seconds' wait per run.

[^isomorphic]: Last year Spike Brehm built a simple project on GitHub that shows how an isomorphic application works in practice. Check it out.

New competitors are already knocking...

Our choice here is actually far from final. A new wave of alternative solutions is already knocking at the door, and it's starting to win over the community's more reactionary (or should we say hipster?) developers.

Webpack

One of the most highly rated tools in this space seems to be Webpack. Without straying too far from Browserify's core concepts, Webpack takes a different stance on a few topics, well described by Browserify's own author in this article, which sparked considerable interest and mixed reactions in the JavaScript community. One to watch.

jspm

Far bolder is the jspm project, which has been getting a lot of attention in recent weeks and whose distinctive features are:

  • installing dependencies from Node.js, Bower and GitHub alike;
  • supporting every library, by implementing all available module-loading mechanisms (globals, AMD, CommonJS);
  • letting you write your application in ECMAScript 6, import directives included;
  • last, and truly innovative when combined with the others: no bundling step required.

How is all this possible? Pre-processing ECMAScript 6 files, parsing import directives and downloading dependencies all happen at runtime, on the client side: a move that, while it slows down the development environment, greatly simplifies bootstrapping a new application, one of the most painful steps for a JavaScript newcomer.

With jspm, in production you can fall back on the classic bundle, or, the project's second highly innovative feature, use an HTTP/2 CDN which, through a mechanism called dependency cache, eliminates the latency and overhead of progressive downloading. We'll try to cover jspm in more depth soon: it's worth it, if only to learn from it :)

Coming next: CSS preprocessors

This post was a tough one :) In the next one we'll take a short break from the JavaScript world to tackle the other half of front-end: stylesheets. Drawing on our experience with Sass, we'll take an objective look at the main JavaScript alternatives available (Less.js and Stylus). Is it worth switching preprocessors?

Follow us on Twitter or via RSS feed to stay posted on the next installment!

footer